OpenAI has quietly revealed that API timeout retries are billable events, meaning every failed request that auto-retries still burns tokens and credits. For Indian startups running high-volume AI workflows on unstable networks or shared hosting, this hidden cost is inflating monthly cloud bills by an estimated 18-25% without any usable output.
Executive Briefing • Key Findings
- Core Update: OpenAI confirms timeout retries and partial response truncation are charged at full token rates, not prorated.
- Key Metrics / Specs: Failed requests costing ₹0.42-₹2.10 per 1K tokens (GPT-4o class) even when no response is returned; retries stack charges rapidly under load spikes.
- Access & Availability: Affects all tiers: Pay-as-you-go, Tier 1-5, with the cost visible only in granular usage logs under "retried_requests" starting this billing cycle.