Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Not GP, but my educated guess is that they are running a system with between 4 and 6 MI325 or MI355x or similar AMD GPUs. With the 50k tps as the total figure for all parallel requests. Those cards have a lot of memory for their price, allowing you to push to really high batch sizes while still having a large context size for each request


4x MI300A in a Gigabyte server




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: