Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost.
I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.
I wonder if higher-availability tiers of LLM providers do a similar thing internally.
Nice turn around, does anyone has a benchmark regarding other types of requests (priority vs send twice) other than voice/call? Or the tests already test that?
If you want a controllable and predictable system, host it yourself. APIs will always have outages, delays and breaking changes every so often. That's the price you pay for not doing it properly.
Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost.
I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.
I wonder if higher-availability tiers of LLM providers do a similar thing internally.
Nice turn around, does anyone has a benchmark regarding other types of requests (priority vs send twice) other than voice/call? Or the tests already test that?
Why not send it thrice?
for a tier thats twice the cost i would expect >2x the speed. somewhere 5-10x
e.g. 1.40m would become 0.30s.
do people really pay for these priority plans?
I love this. Simple. Useful. To the point. If AI was used, I can't tell because it is clearly representing the author's beliefs.
agreed, reads like a breath of fresh air, no fluff
If you want a controllable and predictable system, host it yourself. APIs will always have outages, delays and breaking changes every so often. That's the price you pay for not doing it properly.