Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -28,3 +28,15 @@ All patch modules live in `ddtrace/contrib/internal/{name}/`.
This APM reference lists LLM/AI integrations only to help choose comparable
contrib patch modules. For LLMObs-specific architecture, provider extraction,
streaming, and test transport guidance, use the `llmobs-integrations` skill.

## ElevenLabs Agents

The ElevenLabs patch module wraps conversation worker lifecycle and uses
conversation-local WebSocket proxies. It replaces only the SDK's imported
asynchronous module binding; the global websockets module remains unchanged.
The SDK's audio interfaces and application callbacks are preserved. See the LLMObs
implementation guide for per-turn span layout and audio timing eligibility.

Optional raw vad_score events are observed by the same connection-local proxies
before SDK filtering. They do not require new callbacks, patch points, or event
subscriptions. See the LLMObs guide for the bounded clip-relative activity schema.
Original file line number Diff line number Diff line change
Expand Up @@ -216,3 +216,31 @@ In addition to the full checklist in the apm-integrations [Implementation Guide]
- [ ] `tests/llmobs/suitespec.yml` — LLMObs test suite entry
- [ ] Test dependencies match the suite style; include `vcrpy` only when cassette replay is used
- [ ] `docs/index.rst` — add integration to the docs index

## Hosted voice conversations: ElevenLabs Agents

The ElevenLabs integration instruments the synchronous and asynchronous Agents
WebSocket loops. Its conversation-local socket proxies observe successful sends
and received messages without changing the audio interface or event subscriptions.
Provider event IDs identify turns; text equality and optional completion events do
not. Keep a bounded window of open turns for late interrupted transcripts and flush
it on worker exit. Audio payloads belong only to LLM messages, with a shared encoded
budget for input and output. Direct sibling workflow phases retain approximate
transport-based timing and explicitly annotate metadata.ttfa_eligible = false.
Do not infer TTFA or token usage from these estimates. Client tool execution uses
the SDK's actual handler, with explicit parenting and fail-open instrumentation.

Regression fixtures preserve event order and shapes with synthetic audio and text.
Cover default/custom audio interfaces, both SDK loops, ordinary event subscriptions,
typed messages, interruption, errors, parent context, concurrent conversations,
missing completion events, and audio limits.

Optional ElevenLabs vad_score events produce user_speech_activity metadata only
on the matching user speech phase. Version 1 carries source=elevenlabs_vad,
timing=estimated, and at most 128 clip-relative integer-millisecond intervals_ms
pairs. An empty list means observed silence; absent metadata keeps phase-based
colors. Input WAV bytes and placement anchors remain unchanged. Scores enter speech
at 0.5 and leave at 0.35. Low scores cover at most 500 ms of input samples, with
stale delivery gaps and unobserved portions conservatively active. State follows
input-buffer transfers and shutdown extension; overflow omits activity metadata.
Never turn these observations into TTFA or alter customer event subscriptions.
30 changes: 30 additions & 0 deletions .riot/requirements/1c1de90.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
annotated-types==0.8.0
anyio==4.15.1
attrs==26.1.0
certifi==2026.7.22
charset-normalizer==3.5.2
coverage[toml]==7.16.2
elevenlabs==2.70.0
h11==0.16.0
httpcore==1.0.9
httpx==0.28.1
hypothesis==6.45.0
idna==3.20
iniconfig==2.3.0
mock==5.2.0
opentracing==2.4.0
packaging==26.3
pluggy==1.6.0
pydantic==2.13.5
pydantic-core==2.46.5
pygments==2.21.0
pytest==9.1.1
pytest-asyncio==1.4.0
pytest-cov==7.1.0
pytest-mock==3.16.0
requests==2.34.2
sortedcontainers==2.4.0
typing-extensions==4.16.0
typing-inspection==0.4.4
urllib3==2.8.0
websockets==17.2
33 changes: 33 additions & 0 deletions .riot/requirements/fe9b763.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
annotated-types==0.8.0
anyio==4.15.1
attrs==26.1.0
backports-asyncio-runner==1.2.0
certifi==2026.7.22
charset-normalizer==3.5.2
coverage[toml]==7.16.2
elevenlabs==2.70.0
exceptiongroup==1.3.1
h11==0.16.0
httpcore==1.0.9
httpx==0.28.1
hypothesis==6.45.0
idna==3.20
iniconfig==2.3.0
mock==5.2.0
opentracing==2.4.0
packaging==26.3
pluggy==1.6.0
pydantic==2.13.5
pydantic-core==2.46.5
pygments==2.21.0
pytest==9.1.1
pytest-asyncio==1.4.0
pytest-cov==7.1.0
pytest-mock==3.16.0
requests==2.34.2
sortedcontainers==2.4.0
tomli==2.4.1
typing-extensions==4.16.0
typing-inspection==0.4.4
urllib3==2.8.0
websockets==16.1.1
1 change: 1 addition & 0 deletions ddtrace/_monkey.py
Original file line number Diff line number Diff line change
Expand Up @@ -109,6 +109,7 @@
"tornado": False,
"trio": True,
"openai": True,
"elevenlabs": True,
"langchain": True,
"anthropic": True,
"crewai": True,
Expand Down
22 changes: 22 additions & 0 deletions ddtrace/contrib/internal/elevenlabs/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
"""
The ElevenLabs integration instruments Python Agents conversations over WebSocket.

Enable it with ddtrace-run or ddtrace.patch(elevenlabs=True) and enable
LLM Observability to collect conversation turns, transcripts, client tools, and audio.
Synchronous and asynchronous conversations with default or custom audio interfaces are
supported with ElevenLabs 2.70.0 and later.

Negotiated mono PCM16 audio is wrapped as WAV. Audio has a shared 4 MiB encoded
budget per response; oversized or unsupported audio is omitted while transcripts remain.
Timing describes audio sent or handed to the SDK, with estimated playback and interruption
truncation. These turns opt out of time-to-first-agent-audio measurements.

For improved user-speaking colors, enable the optional ``vad_score`` client event
in the ElevenLabs agent's Advanced settings under Client Events and save the agent.
Retain the application's other client events. The integration never changes agent
configuration or event subscriptions. Without usable VAD, existing audio playback
and phase-based colors remain available. VAD timing is estimated and does not enable
latency measurements.

Standalone speech-to-text, text-to-speech, and browser transports are not instrumented.
"""
Loading
Loading