Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

But surely it can be scaled up, or is this compression thing something making the approach good only for small models (I haven't read the Deepseek papers (can't allocate time to it))?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: