Skip to content

Tolerate incomplete model listings from OpenAI-compatible endpoints, and bump openai_dive to 1.4.3 - #117

Merged
Marlinski merged 2 commits into
mainfrom
fix/lenient-models-listing
Aug 31, 2026
Merged

Marlinski merged 2 commits into
mainfrom
fix/lenient-models-listing

Conversation

@Marlinski

Copy link
Copy Markdown
Collaborator

Problem

shai auth fails to list models against some OpenAI-compatible endpoints with:

Failed to fetch models: missing field `object`

The cause is deserialization strictness, not the endpoint. openai_dive declares:

pub struct ListModelResponse { pub object: String, pub data: Vec<Model> }
pub struct Model { pub id: String, pub created: Option<u32>, pub object: String, pub owned_by: String }

Three of those fields are mandatory String. Plenty of gateways return a perfectly usable
{"data":[{"id":"vendor/some-model"}, ...]} with no top-level "object":"list", no per-entry
"object":"model", and sometimes no owned_by. serde aborts on the first missing one, so the
whole listing is rejected over metadata shai never reads — it only needs the ids. This breaks
model selection in shai auth and default_model() for those providers.

Fix

New shai_llm::providers::models::list_models_compat, which fetches /models with the client's
own http client (the pattern mistral.rs already used, since Client::get is pub(crate)) and
requires only id per entry:

  • defaults object to "model", leaves owned_by / created empty when absent
  • also accepts a bare top-level array, which some gateways return
  • skips entries without a usable id instead of failing the whole listing
  • reports non-2xx with the status and a truncated body, instead of a confusing parse error

Wired into the four providers that went through the strict listing: openai,
openai_compatible, ovhcloud and ollama. mistral, openrouter and anthropic already
had their own conversions and are untouched.

Note that bumping openai_dive does not fix this on its own — 1.4.3 still declares those
fields as mandatory.

Also: bump openai_dive 1.3.1 -> 1.4.3

Second commit, across all four crates. Breaking changes handled:

Change in 1.4.x Handling
ChatMessage::Assistant and both DeltaChatMessage variants gained reasoning set in all initializers
shared::Usage lost input_tokens, output_tokens and their _details dropped from the chat Usage literals — every site set them to None
Responses API got its own ResponseUsage carrying those fields response/formatter.rs builds a ResponseUsage
ResponseObject gained top_logprobs; items::FunctionToolCall made id and status optional updated

The ResponseUsage split also fixes a latent spec bug: /v1/responses was reporting
chat-style prompt_tokens / completion_tokens inside a Responses payload, where the spec
calls for input_tokens / output_tokens.

Providers are split on how they report reasoning — some send reasoning, some
reasoning_content. On 1.3.1 the former did not exist in the struct at all, so it was silently
discarded. extract_think_content now normalizes reasoning onto reasoning_content, which is
what the agent and the TUI read, so reasoning from those models is displayed rather than
dropped. The existing <think> extraction still takes precedence.

Testing

  • 5 unit tests on the lenient parser: missing object fields, a standard OpenAI listing, a bare
    array, entries without an id, and an unparsable body. No network or credentials needed.
  • 5 tests on what the bump could regress silently. 1.4.x switched several enums from
    rename_all = "lowercase" to "snake_case", so role tags, tool types and tool-choice values
    are asserted to serialize unchanged (they do — every variant shai uses is a single word),
    alongside the reasoning normalization and <think> extraction.
  • Verified against a live OpenAI-compatible gateway whose listing omits object: model listing,
    chat with function calling, tool-call parsing and token usage all work, for FunctionCall,
    FunctionCallRequired and StructuredOutput alike. Also exercised end to end through the
    release binary in headless mode — listing, chat, tool call, tool execution, reasoning display.
  • No test regressions: the same 36 pre-existing failures before and after (live tests needing
    credentials, plus an unrelated ringbuffer panic in fc::tests::test_clear_history).
  • cargo build --release clean; the one warning is pre-existing and unrelated
    (shell/pty.rs:193, a rustc function-cast lint).

Out of scope, noticed along the way

  • fc::tests::test_clear_history panics inside ringbuffer-0.16 and fails on main too.
  • LlmClient::default_model() returns Ok("") when SHAI_MODEL is set but empty, instead of
    falling back to the provider.

openai_dive deserializes GET /models into ListModelResponse, where the
top-level `object` field and each Model's `object` / `owned_by` are
mandatory Strings. Several openai-compatible gateways omit them, so the
whole listing failed with an opaque "missing field `object`" even though
the payload contained perfectly usable model ids. This broke `shai auth`
model selection and default_model() for those endpoints.

Add providers::models::list_models_compat, which fetches the listing with
the client's own http client and only requires `id` per entry, defaulting
object to "model" and leaving owned_by/created empty when absent. It also
accepts a bare top-level array, skips entries without a usable id, and
reports non-2xx responses with the status and a truncated body instead of
a parse error.

Used by the four providers that went through the strict listing: openai,
openai_compatible, ovhcloud and ollama. Upgrading openai_dive would not
help here: 1.4.3 still declares those fields as mandatory.
1.3.1 -> 1.4.3 across shai-llm, shai-core, shai-cli and shai-http.

Breaking changes handled:

- ChatMessage::Assistant, DeltaChatMessage::Assistant and
  DeltaChatMessage::Untagged gained a `reasoning` field alongside the
  existing `reasoning_content`. All initializers now set it.

- shared::Usage lost input_tokens, input_tokens_details, output_tokens and
  output_tokens_details; the Responses API got its own ResponseUsage type
  carrying them instead. Every site we had set those four to None, so they
  are simply dropped from the chat Usage literals. The Responses API
  formatter now builds a ResponseUsage, which also makes its emitted usage
  object spec-correct: it previously reported chat-style prompt_tokens /
  completion_tokens inside a Responses payload.

- ResponseObject gained top_logprobs, and items::FunctionToolCall made `id`
  and `status` optional.

Providers are split on how they report reasoning: some use `reasoning`,
some `reasoning_content`. extract_think_content now normalizes the former
onto the latter, which is what the agent and the TUI read, so reasoning
from those models is displayed instead of silently dropped. The existing
<think> extraction still takes precedence.

Added tests covering the parts of the upgrade that could regress silently:
1.4.x switched several enums from rename_all = "lowercase" to "snake_case",
so role tags, tool types and tool-choice values are asserted to serialize
unchanged, along with the reasoning normalization and <think> extraction.

Verified against a live openai-compatible endpoint: model listing, chat
with function calling, tool-call parsing and token usage all work, for
FunctionCall, FunctionCallRequired and StructuredOutput alike. No test
regressions: the same 36 pre-existing failures (live tests needing
credentials, plus an unrelated ringbuffer panic) before and after.
@Marlinski
Marlinski merged commit 707c619 into main Aug 31, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant