A team is preparing a pipeline for documents of exactly 600 words each.
At word level, their vocabulary covers 95% of the word occurrences in these documents. The remaining 5% fall outside it.
They switch to a subword tokenizer, which behaves like this on their data:
The trainer then pads every document out to a fixed length of 768 tokens. No document here is long enough to be truncated.
What percentage of the 768 slots in one document is padding?
Round your answer to 2 decimal places, given as a percentage — for example 7.25, not 0.0725.