A lab has trained three language models on the same kind of text, each one with its data and compute grown alongside its parameter count, so none of them is starved or undertrained. Only the scale differs.
| parameters | test loss |
|---|---|
| million | |
| billion | |
| billion |
The lab is now budgeting for a 1-trillion-parameter model, trained the same way, and wants a forecast of its test loss before spending anything.
Extending the regularity in the table, what loss should they forecast?
Select all that apply.