Skip to content

Opt-in automatic retry honoring the Retry-After header on 429 #177

Description

@teionarr

The SDK does no automatic retry, and the keyed-client 429 path discards the server's documented Retry-After header — so every production user hand-rolls their own backoff.

Impact (why it's worth prioritizing)

  • Lands on the primary use case. Agentic / deep-research workloads fan out many searches concurrently — exactly what trips per-minute limits — so 429s hit the workloads Tavily is built for, precisely when they're under load.
  • Wasted latency by design. The server returns the exact wait via Retry-After, but the SDK drops it, so callers must guess backoff — too long slows every rate-limited call; too short causes repeat 429s and risks escalation.
  • Silent quality degradation, not just failures. In pipelines that swallow retriever errors to [] (e.g. gpt-researcher), a dropped 429 becomes empty results → the LLM fabricates, rather than a clean, recoverable error.
  • Duplicated across the whole user base. Everyone re-implements the same retry + backoff + Retry-After parsing — the ~60 lines feat(errors): expose Retry-After header on UsageLimitExceededError #166 already wrote.

Current behavior

tavily/tavily.py:134 (async tavily/async_tavily.py:162): a 429 raises UsageLimitExceededError immediately; the retry-after header is dropped, and there's no max_retries/backoff in either client.

Why this belongs in the SDK

Peer vendor SDKs ship this by default — OpenAI & Anthropic default to automatic retries honoring retry-after; Stripe retries with idempotency keys. The rate-limits docs already recommend honoring retry-after, and #166 adds the parsing — this is the natural next step: acting on it, opt-in.

Proposal (opt-in, default-off — no behavior change unless enabled)

  • TavilyClient(..., max_retries=0) / AsyncTavilyClient(...)
  • Retry 429, 502, 503, 504 + connection errors; never 400 / 401 / 403 / 432 / 433
  • Honor Retry-After (feat(errors): expose Retry-After header on UsageLimitExceededError #166's _parse_retry_after already handles integer seconds + HTTP-date); otherwise exponential backoff + full jitter, capped; respect the caller's timeout
  • Idempotency-aware: safe on /search, /extract, /map, get_research; retry /crawl and /research on 429 only (not ambiguous read/connect timeouts → avoid duplicate billable jobs); never retry mid-stream

Builds directly on #166 — happy to open a PR (sync + async + tests) if you'd welcome it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions