You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The SDK does no automatic retry, and the keyed-client 429 path discards the server's documented Retry-After header — so every production user hand-rolls their own backoff.
Impact (why it's worth prioritizing)
Lands on the primary use case. Agentic / deep-research workloads fan out many searches concurrently — exactly what trips per-minute limits — so 429s hit the workloads Tavily is built for, precisely when they're under load.
Wasted latency by design. The server returns the exact wait via Retry-After, but the SDK drops it, so callers must guess backoff — too long slows every rate-limited call; too short causes repeat 429s and risks escalation.
Silent quality degradation, not just failures. In pipelines that swallow retriever errors to [] (e.g. gpt-researcher), a dropped 429 becomes empty results → the LLM fabricates, rather than a clean, recoverable error.
tavily/tavily.py:134 (async tavily/async_tavily.py:162): a 429 raises UsageLimitExceededError immediately; the retry-after header is dropped, and there's no max_retries/backoff in either client.
Why this belongs in the SDK
Peer vendor SDKs ship this by default — OpenAI & Anthropic default to automatic retries honoring retry-after; Stripe retries with idempotency keys. The rate-limits docs already recommend honoring retry-after, and #166 adds the parsing — this is the natural next step: acting on it, opt-in.
Proposal (opt-in, default-off — no behavior change unless enabled)
The SDK does no automatic retry, and the keyed-client 429 path discards the server's documented
Retry-Afterheader — so every production user hand-rolls their own backoff.Impact (why it's worth prioritizing)
Retry-After, but the SDK drops it, so callers must guess backoff — too long slows every rate-limited call; too short causes repeat 429s and risks escalation.[](e.g.gpt-researcher), a dropped 429 becomes empty results → the LLM fabricates, rather than a clean, recoverable error.Retry-Afterparsing — the ~60 lines feat(errors): expose Retry-After header on UsageLimitExceededError #166 already wrote.Current behavior
tavily/tavily.py:134(asynctavily/async_tavily.py:162): a 429 raisesUsageLimitExceededErrorimmediately; theretry-afterheader is dropped, and there's nomax_retries/backoff in either client.Why this belongs in the SDK
Peer vendor SDKs ship this by default — OpenAI & Anthropic default to automatic retries honoring
retry-after; Stripe retries with idempotency keys. The rate-limits docs already recommend honoringretry-after, and #166 adds the parsing — this is the natural next step: acting on it, opt-in.Proposal (opt-in, default-off — no behavior change unless enabled)
TavilyClient(..., max_retries=0)/AsyncTavilyClient(...)429, 502, 503, 504+ connection errors; never400 / 401 / 403 / 432 / 433Retry-After(feat(errors): expose Retry-After header on UsageLimitExceededError #166's_parse_retry_afteralready handles integer seconds + HTTP-date); otherwise exponential backoff + full jitter, capped; respect the caller'stimeout/search,/extract,/map,get_research; retry/crawland/researchon 429 only (not ambiguous read/connect timeouts → avoid duplicate billable jobs); never retry mid-streamBuilds directly on #166 — happy to open a PR (sync + async + tests) if you'd welcome it.