DeepMind 70B 参数计算最优语言模型,1.4 万亿 token 训练,超越 Gopher 280B 与 GPT-3。
提出 Chinchilla 缩放定律:数据与算力需同步增长,70B 模型可达更优性能。