A text classifier uses an LSTM over 120-dimensional embeddings, with a hidden state of 180 units.
Inside both an LSTM and a GRU, every gate and every candidate is computed the same way — from the current token and the previous hidden state:
each with its own , its own and its own bias . ( is the squashing function and holds no parameters.) Call one of these a block.
Nothing else in either cell holds parameters. So for input size and hidden size , one block holds numbers in , numbers in , and biases.
You now replace the LSTM with a GRU and spend the entire saving on a wider hidden state: the GRU must use exactly the same total number of parameters as the LSTM it replaces.
What hidden size does the GRU get?
The answer is a whole number. (The embedding table and the output layer sit outside this count and are unchanged.)