Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Good time to buy then, I don’t understand how stupid some traders can be.

Likely a "how solid is the technical moat" evaluation - this could be a one-off or could be that there are an avalanche of advancements to continue along the efficiency side of the process.

Given the style and hype of logic in the AI space, I fully believe resources are not well allocated in compute and _actual_ thinking as to how they are spent.

Deepseek's apparent 10x more efficient per inference token... implies a lot of other hardware meets the general use-case. We also know that reasoning should be about 10W for human speed-of-thought... maybe another 1-2 orders of power efficiency.

"Pre-Training: Towards Ultimate Training Efficiency

We design an FP8 mixed precision training framework and, for the first time, validate the feasibility and effectiveness of FP8 training on an extremely large-scale model. Through co-design of algorithms, frameworks, and hardware, we overcome the communication bottleneck in cross-node MoE training, nearly achieving full computation-communication overlap. This significantly enhances our training efficiency and reduces the training costs, enabling us to further scale up the model size without additional overhead. At an economical cost of only 2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model. The subsequent training stages after pre-training require only 0.1M GPU hours." [1]

[1] https://huggingface.co/deepseek-ai/DeepSeek-V3



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: