Skip to content

feat(llmobs): standardize agent manifest across integrations - #20731

Draft
mz1119 wants to merge 3 commits into
mainfrom
max.zhang/llmobs-standardize-agent-manifest
Draft

mz1119 wants to merge 3 commits into
mainfrom
max.zhang/llmobs-standardize-agent-manifest

Conversation

@mz1119

@mz1119 mz1119 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Description

Standardizes the agent_manifest emitted on agent spans across CrewAI, Google ADK, LangGraph, OpenAI Agents and Claude Agent SDK, so every integration reports the same keys (AgentManifest) in the same shapes, and instructions, tools and model are populated wherever the framework declares them.

Shared helpers in _integrations/agent_manifest.py (build_agent_manifest, instruction_fields, normalize_tool, filter_model_settings, config_value, as_str) generalize what pydantic_ai and the manual path already did: section-isolated building, an allowlist for model settings, a JSON-native result, and callables recorded by name rather than called or str()-ed. pydantic_ai moves onto them; its output only changes in key order. A failing section now emits an agent_manifest.section_error telemetry count tagged with integration and section.

Tool parameters are normalized to {param: {type, required?}} with JSON Schema type names. Optional parameters (anyOf, list-valued type, $ref) keep a type instead of being dropped.

Per integration:

  • OpenAI Agents: model_settings is allowlisted (it previously shipped extra_headers, extra_body, extra_query, extra_args). Hosted tools ship only declared config (no MCP headers, authorization or URL, no client objects, no web search user location). MCP servers left at their default name (which embeds the URL) report the class name. Callable instructions and stored prompts go to extra_instructions. Adds capabilities, data_contracts and agent_settings.
  • Google ADK: string models are no longer dropped. Callable instructions ship by name instead of as a memory address. description maps to handoff_description. Adds system_prompts, model_settings (from generate_content_config), tool parameters for every callable tool (bound methods and partials included), data_contracts, handoffs (sub_agents), guardrails (before-model/tool callbacks) and agent_settings. Removes model_configuration (pydantic's own config) and session_management (per-run values).
  • CrewAI: instructions is goal plus backstory, preferring the pre-interpolation text when CrewAI kept it. Tool parameters come from args_schema. max_iter and code-execution settings move to agent_settings. CrewAI's injected stop words are not reported.
  • LangGraph: create_react_agent(model="gpt-4o") no longer raises. Tool parameters come from tool_call_schema, so injected args are excluded. The cached manifest is no longer mutated by a run's config; recursion_limit moves to agent_settings, preferring the one declared with with_config(). dependencies (run input keys) is removed.
  • Claude Agent SDK: adds instructions (system_prompt, including presets), capabilities (MCP servers, without per-run status), handoffs (agents), guardrails and agent_settings, for both query() and ClaudeSDKClient sessions. Removes dependencies/max_iterations.

Every key web-ui reads (model_provider, handoff_description, CrewAI's {allow_delegation} handoffs) is kept. model_configuration was only a fallback for a missing model_settings, which ADK now reports.

Supersedes #19533.

Testing

  • Each integration's expected manifests updated against the real frameworks; CI covers the full library version matrix.
  • tests/llmobs/test_agent_manifest_integrations.py checks that every integration emits only schema keys, round-trips through JSON, and is identical across rebuilds, and that no declared callable is invoked. It covers optional parameters, ADK callable tools and string models, LangGraph injected args, model strings without a colon and cache isolation, OpenAI secret-bearing settings, tools and MCP names, and Claude MCP fallback.
  • A ClaudeSDKClient(options=...) test covers options-derived fields on client sessions.

Risks

Manifest keys change shape for existing integrations (listed above and in the upgrade release note). Comparing agent versions recorded before and after upgrading, or across services on different tracer versions, shows differences caused by the tracer. Reverting this PR would also revert the OpenAI credential allowlist.

Additional Notes

The agent_manifest.section_error telemetry metric may need registering on the telemetry intake before it is visible.

🤖 Generated with Claude Code

Move CrewAI, Google ADK, LangGraph, OpenAI Agents and Claude Agent SDK onto
shared manifest helpers so every integration emits only AgentManifest keys,
with instructions, tools and model populated where the framework declares them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against main using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

.claude/skills/llmobs-integrations/SKILL.md                             @DataDog/python-guild
ddtrace/contrib/internal/claude_agent_sdk/patch.py                      @DataDog/ml-observability
ddtrace/llmobs/_integrations/agent_manifest.py                          @DataDog/ml-observability
ddtrace/llmobs/_integrations/claude_agent_sdk.py                        @DataDog/ml-observability
ddtrace/llmobs/_integrations/crewai.py                                  @DataDog/ml-observability
ddtrace/llmobs/_integrations/google_adk.py                              @DataDog/ml-observability
ddtrace/llmobs/_integrations/langgraph.py                               @DataDog/ml-observability
ddtrace/llmobs/_integrations/openai_agents.py                           @DataDog/ml-observability
ddtrace/llmobs/_integrations/pydantic_ai.py                             @DataDog/ml-observability
ddtrace/llmobs/_telemetry.py                                            @DataDog/ml-observability
ddtrace/llmobs/types.py                                                 @DataDog/ml-observability
releasenotes/notes/llmobs-standardize-agent-manifest-df86615c800061b2.yaml  @DataDog/apm-python
tests/contrib/claude_agent_sdk/test_claude_agent_sdk_llmobs.py          @DataDog/ml-observability
tests/contrib/claude_agent_sdk/utils.py                                 @DataDog/ml-observability
tests/contrib/crewai/test_crewai_llmobs.py                              @DataDog/ml-observability
tests/contrib/google_adk/test_google_adk_llmobs.py                      @DataDog/ml-observability
tests/contrib/langgraph/test_langgraph_llmobs.py                        @DataDog/ml-observability
tests/contrib/openai_agents/test_openai_agents_llmobs.py                @DataDog/ml-observability
tests/llmobs/test_agent_manifest_integrations.py                        @DataDog/ml-observability

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

📈 Existing violations got worse

3 pre-existing violation(s) increased in severity (e.g. their target became more depended-on, or got pulled into an import cycle), though the edge itself isn't new:

ddtrace.contrib.internal.openai._realtime -×-> ddtrace.llmobs.types  (contrib -> product:llmobs, score=38, +2 vs base)
ddtrace.contrib.internal.claude_agent_sdk._streaming -×-> ddtrace.llmobs.types  (contrib -> product:llmobs, score=38, +2 vs base)
ddtrace.contrib.internal.vllm.extractors -×-> ddtrace.llmobs.types  (contrib -> product:llmobs, score=38, +2 vs base)

⚠️ Existing dependency direction violations

There are 201 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 201 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.profiling.collector.pytorch -×-> ddtrace.trace  (product:profiling -> product:tracing, score=130)
ddtrace.llmobs._integrations.llama_index -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.llmobs._integrations.openai_agents -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.profiling.scheduler -×-> ddtrace.trace  (product:profiling -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@datadog-prod-us1-6

datadog-prod-us1-6 Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Pipelines  Tests

✨ Unblock PR with BitsAI

❌ Errors

Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 7 Pipeline jobs failed

DataDog/apm-reliability/dd-trace-py | llmobs/google_adk 1/4 — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

DataDog/apm-reliability/dd-trace-py | llmobs/google_adk 3/4 — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

DataDog/apm-reliability/dd-trace-py | llmobs/llmobs 1/5 — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

View all 7 failed jobs.

ℹ️ Info

No other issues found (see more)

🧪 All tests passed
❄️ No new flaky tests detected

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 4c29585 | Docs | View more details | Give us feedback!

mz1119 and others added 2 commits October 5, 2026 13:22
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep optional tool parameters, read ClaudeSDKClient options, stop reporting
default MCP server names and web search user location, exclude LangGraph
injected tool args, report every ADK callable tool, prefer CrewAI's declared
goal and backstory, count manifest section failures in telemetry, and
correct the release note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mz1119

mz1119 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-06T21:15:32.315625Z 4c29585 Manual request
🔒 Security Review ✅ Completed 2026-10-06T21:15:53.834061Z 4c29585 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Can't wait for the next one!

Reviewed commit: 4c295855a1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

🛡️ Codex Security Review · Automatically triggered

Security review completed. No security issues were found in this pull request.

Reviewed commit: 4c295855a1

View security finding report

Only the user who started this review can view the report in Codex.

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant