A note-taking app is adding a Summarise this meeting button. Each press sends the model a prompt of 1,800 tokens (the transcript plus instructions) and gets back a reply of 350 tokens.
The team is choosing between two ways to serve it.
Option A: a hosted API. Another company runs the model, and you pay for every token:
| token | price |
|---|---|
| prompt tokens (read by the model) | $2.50 per million |
| reply tokens (written by the model) | $10.00 per million |
Option B: your own GPU. You rent a dedicated GPU server and run a model of the same quality on it:
| cost | amount |
|---|---|
| GPU rental | $1.85 per hour, billed for every hour of the day whether requests arrive or not |
| upkeep (an engineer's time keeping it patched and running) | $1,080 per 30-day month |
One GPU easily keeps up with every request volume in this question, and nothing else costs money in either option.
At how many requests per day do the two options cost exactly the same? Round to the nearest whole number.