Skip to content

Commit c21ba2b

Browse files
authored
Setup, end to end: pick a Bot, connect a model, prove it answers (#451)
* Give a signed-in ChatGPT plan the Codex model, and the store rather than a token A plan token is a bearer for chatgpt.com/backend-api/codex, and langchain-openai pins that address on purpose, so it cannot be reached by pointing OPENAI_BASE_URL at it with the token as a key. The harness now picks the Codex chat model when a plan is present, so the default Bot can actually answer on a subscription. Carrying the access token alone was wrong: it expires within the hour and nothing can renew it, which would give a Bot that works in the morning and fails after lunch with an auth error nobody could account for. The sign-in now hands back the vendor's whole store, the store is written beside the .env as an owner-only file, and compose bind-mounts it read-write so the renewals the provider makes outlast the container. The .env gets a path, never the credential. The file is written even when no plan was chosen, because a bind mount with no source does not fail, it silently creates a directory in its place. * Empty the plan token this app has stopped writing A machine that ran the version before this one has a ChatGPT plan token sitting in its .env that nothing reads any more. The writer preserves lines it does not own, so it would stay there indefinitely. Clearing it costs one line and is the same reasoning the other model keys are emptied for. * End the install with a question the Bot has to answer Every step before this proves that something started, which is not the same as proving the choices work. A refused key, a lapsed plan or a model the account cannot use all give a stack that comes up clean and a Bot that cannot answer, and handing over at that point means somebody finds out later, inside the product, with no idea which of their answers caused it. So the wizard now ends on a question with one checkable answer, and the handover waits for it. This screen also owns the worst message in the product. Measured against a deliberately invalid key: the stream opens, says RUN_STARTED, says STEP_STARTED and then simply stops, with no error event at all, because the framework caught its own exception and logged it. The whole 401 lives in the container's log and nowhere else. So a run that produces no text is a failure here rather than an empty answer, the sentence shown is OpenBot's own and names the choice to change, and the harness's log is fetched to fill the developer half, since otherwise there would be no developer half to show. Two live tests are kept and ignored by default. The fixtures are transcriptions of a real stream, and a vendor changing the events they emit should break something. * Keep credentials in the machine's own store, not in the .env The .env is a settings file, and a settings file is something somebody opens, reads out to support or pastes into a chat. A model key, a plan token and the tokens these services prove themselves to each other with are not settings. They now go to the credential store each platform actually has: the login Keychain on macOS through security, DPAPI on Windows through ProtectedData encrypting to the signed-in user, and an owner-only file on Linux, which is said out loud rather than dressed up, because no desktop Linux install can be assumed to run a Secret Service daemon and refusing to save a credential because gnome-keyring is missing would fail more people than it protects. The value never goes on a command line on any of them. ps is readable by every process the person runs, so the Keychain and DPAPI paths both write over stdin. From the store the credentials travel to the containers and the host processes as environment. Compose resolves an interpolation from its own environment before it reads the .env, so a secret reaches exactly the services that declare it and is written down nowhere. Verified against a real deployment: a .env with no token, the token in the environment, and the container holding it. The writer also purges what it moved. Without that, every machine that ran an earlier version would keep its old plaintext copy exactly where it was and the change would have bought nothing for anybody who already had OpenBot. * Give the secondary buttons the class the stylesheet actually defines Three buttons asked for `secondary`, which no rule matches, so Stop OpenBot has been rendering identically to Show OpenBot and the new last screen offered two equally weighted actions. The stylesheet's quiet button is what they meant. Found by looking at the screen rather than at the markup, which is the only way a missing class shows up: nothing errors, the button just draws as the primary. * Store a credential through the Keychain itself, not through the security command The `security` command takes its password through a prompt whose buffer is 128 bytes, and anything longer is cut off with no error and an exit status of zero. Probed a length at a time: 128 stores 128, 129 stores 128, 200 stores 128. An OpenAI project key is 164 characters. Every one of them was being saved truncated and read back truncated on the next run, while the run that saved it worked fine, because the value it used came straight from the window. The next launch would have been the broken one, with a key nobody had changed. No flag raises that buffer, and the only ways past the prompt put the credential on a command line where ps can read it. The framework has neither limit, and the Windows and Linux paths never had one, so this is one platform's dependency rather than a cross-platform crate and twenty transitive packages. Round-tripped against the real Keychain at 128, 129, 164, 256 and 512. * Mint the generated secrets once per deployment, not once per start KEY_ENCRYPTION_KEY is what every secret the server stores goes through, and encrypt-sso-config.ts names the symptom itself: a changed key leaves stored configuration unreadable and sign-in broken until it is registered again. The shell minted a new one on every press of Start, so a button labelled Start silently orphaned everything the previous run had encrypted. The rest of the generated secrets point the same way for a smaller reason: a Bot's computer is a container that outlives a restart holding the old COMPUTER_TOKEN, so rotating buys nothing and can only strand it. Two installs still do not share a key. A machine with nothing stored generates, which is what a first run is. Verified across a real Stop and Start: the token the second run handed its containers is the one the first run minted. * Stop the Bot that was picked, and stop refusing to start because of it Two halves of one dead end, both found by pressing the buttons. Stop left agent-harness running. Compose only acts on a profiled service when the profile is named, so the one container the person actually chose stayed up on their laptop after they had stopped the app, still holding its port. Then the next Start refused: "something is already listening on port 4206, which OpenBot uses for the Bot you picked" — about a container OpenBot itself had started, which the person never saw and could not find. There was no way forward from that screen. A port this deployment already publishes is not a stranger on the port, and compose up reuses what is there, so the check now skips our own and keeps its teeth for somebody else's. Host processes are reclaimed for the same reason: a start that got as far as spawning the server and then stopped left it holding 3001, and they are identified by working directory, so anything stopped belongs to this deployment and no other. Verified in the window: Stop leaves nothing running, and a Start with our own container on the port goes straight through. * Make the ChatGPT sign-in work, and say what went wrong when it does not Four things, all found by signing in through the window rather than testing the pieces. The login program is rendered by substituting placeholders, and renaming the marker to OPENBOT_CHATGPT_STORE made the marker contain the STORE placeholder. Rendering rewrote the program's own print line into a syntax error, so the container died before printing anything and the window said "the sign-in never offered a link to open" — a sentence with no relation to its cause. The placeholders are underscored now and the test reads the rendered program rather than the template, which is where it broke. That failure also had no technical half, so there was nothing to diagnose it with. It carries the container's output now, with the store line stripped. The window rendered a failure with String(error), which for a two-fold problem prints [object Object]. There is one failure component now and every screen uses it. And the last screen swallowed its own failure entirely: it recorded that something went wrong and threw the problem away, believing the screen around it would show the sentence. Nothing did. A plan that could not answer produced a "Change the model" button and no words at all — the exact silence that screen exists to replace. Proven in the window: consent in the browser, the store written owner-only with its refresh token, mounted into the harness, and the Bot answering 17 x 23 = 391 through the Codex model with no OpenAI key in the container, the .env or the Keychain. * Give the window an Edit menu, and let a plan pick the Bot that can spend it Two defects that only the Anthropic path could show, both on the last screens a person sees. macOS routes the clipboard shortcuts through the menu bar, and this window had no Edit menu, so it had no Paste. Typing into the code field worked and pasting did nothing — on the one screen whose own instruction is "paste the code it shows you". Everybody signing in to a Claude plan would have reached that field, pressed the shortcut they have used all their life, and had nothing happen. Then, with the code in: a plan is not a key, and only one Bot speaks each vendor's subscription. Signing in to Claude and keeping the default Bot gave a stack that came up clean and a Bot whose log read "Missing credentials. Please pass an `api_key`". The person had answered both screens correctly and had no way to know which answer to change. The plan now re-points the Bot, and the model screen says which Bot that will be while there is still a screen to say it on. Nobody is asked to know that a subscription constrains the framework. Proven in the window on both plans: ChatGPT answers 17 x 23 = 391 through Codex with no OpenAI key present, and Claude answers 391 on the Claude Agent SDK with no Anthropic key present. The refusal of a stale code and the failure of a Bot that cannot answer both render with a plain sentence and the container's own log behind a disclosure. * Make the CopilotKit sign-in work, which took four shape mismatches to find Signing in from the window had never been run end to end. It failed four times in a row, each time silently or with a message that named nothing, and each fix was only findable because the failure started carrying what actually came back. The session is called `cliToken`, not `token`, so the very first exchange failed with "error decoding response body" and no way to tell which field or which endpoint. Failures here now carry the response, and that answered it in seconds. Project ids are numbers. Requiring a string dropped every project, and the screen said "That account has no projects yet" to somebody with ten of them. An empty list and an unreadable one are told apart now, because one of them is a lie a person cannot argue with. The keys endpoint declares `project_id: z.number()` with no coercion, so the string "7" came back HTTP 400 VALIDATION_ERROR on the last step of the flow. Read off `api-keys-routes.ts` rather than guessed. And the project tiles rendered as blank white rectangles: `button` sets a white colour, `.tile` overrode the background to white and not the colour, and the provider rows escaped it only because they are labels. Nothing errored. The screen asked somebody to choose between six empty boxes. A shown body is masked, because the one that diagnosed the first bug also carried a live session token, and the shape is what a developer needs from it. Proven in the window: sign in, choose an organisation, ten projects listed by name, one picked, a key created, "Connected to CopilotKit". No key typed. * Stop a model name outliving the answer that chose it The compatible row is the only one that names a model, and switching away from it kept the name. Answering with an OpenAI key after using a local endpoint left BOT_MODEL=local-model, so the Bot asked OpenAI for a model only that person's own server has, and the last screen said "That account cannot use the model that was chosen" about a model this run never chose. Exactly the failure the key-clearing exists for, with one key missed. It is removed rather than emptied so the compose default applies, and taken out of the file as well, because the writer keeps lines it does not own and that is what let it survive. Found on a full pass through the window, in the first path. * Say why a conversation has no messages, instead of showing a blank window Clicking a conversation in the rail drew the coworker's name and then nothing at all. The rail comes from OpenBot's own database, so a channel is listed whatever the history store says; the messages live in the Intelligence project, and pointing a deployment at a different project leaves the platform answering THREAD_NOT_FOUND. That 404 is deliberately read as "no history" and must stay that way: a thread id is minted before the thread exists, so a brand-new conversation 404s as its normal opening move. Widening it would tell somebody their conversation was gone and invite them to start it over. The two cases are told apart by a fact the app already stores. lastMessageAt is set only once something has been said, so a conversation with none is genuinely new and silence is correct, while one that has been spoken in and comes back empty has a history this deployment cannot reach. That one now says so, in the notice slot beside the existing explanations for a deleted coworker and for turns that could not be parsed. The channel DTO carries lastMessageAt for it, which the shape tests pin, plus a new test that the date leaves as a string and leaves at all. * Install the container engine, instead of telling somebody to go and get one Setup ended at "Install Podman Desktop or Docker Desktop first" on any machine that had neither, which is every machine this app is for. The step existed in the enum with nothing behind it and the screens had been reworded to stop promising it. So the whole install stopped at a download page. It installs one now, and the second half is the half that gets forgotten: Podman ships no Compose implementation, so a machine with a freshly installed Podman still cannot raise the stack and answers with seven errors naming docker-compose. Both are fetched, each pinned to the digest of the release it was tested against and refused if it does not match, because these are files this app then executes. Only what is missing is added: an engine somebody already has is theirs, and a Compose that already answers is left alone. Windows installs unattended. macOS and Linux each raise one authorization prompt, which is the platform's own and is not something to route around: the package writes to /opt/podman, and on Linux Podman is a set of binaries wired to the distribution's paths rather than one file to download. Two things were needed to make the result usable in the session that installed it. The MSI extends the USER's PATH, and this process was started with the old one, so podman could not be run for the rest of the run: every engine command now names a resolved path, found on PATH first and in the installers' own locations second. And the Compose provider is put in front of the child's PATH rather than written into containers.conf, which belongs to whoever else may have configured it. Measured on Windows Server 2022, which is also where the recovery came from: a Podman removed by deleting its folder leaves the registration behind, so /i becomes a repair with no source and stops with 1603. That case uninstalls and installs cleanly instead of reporting a failure somebody cannot act on. Image references now come from the release's manifest wherever the shell runs a container itself, not only where Compose does. This is the same bug a third time: first the names were built from the ids and matched nothing published, then the version stopped being appended so an engine read the bare name as :latest, and now openbot-agent-langgraph-agui:v0.0.8 was resolved to docker.io/library/... and the person was told access was denied, which reads as a credentials problem for a repository that was never pushed. No reference is built here at all any more, and an image this release does not include is named as that. Both plan sign-ins set the engine up rather than refusing. They run in a container, and "No container engine is answering, so the sign-in cannot run" named an obstacle and no way past it, on a screen whose whole purpose is to put one there. One function does it for Start and for both of them. A setup step that stops now carries both registers. podman machine init failing is exactly the case the two-part failure was written for, and it was the last place still putting an engine's own words in front of somebody as the headline. * Let an endpoint that needs no key be connected The compatible row refused to continue without an API key, and its own summary names Ollama and vLLM. Neither has one. So the two examples the screen offers by name were the two it would not accept, and the way out was to invent a key and hope the endpoint ignored it. An address and a model name are what that row needs. The Rust side already treated the key as optional and writes OPENAI_API_KEY only when it is given, so the refusal lived entirely in the screen. The field says what it is now rather than leaving somebody to find out by being stuck. A failure also belongs to the row that produced it. A refused OpenAI sign-in stayed on screen after switching to the endpoint row, underneath the address just typed, where it read as a complaint about that address. Found by driving the screen on Windows. * Let a Bot answer from an endpoint that needs no key Making the compatible-endpoint row accept a blank key fixed one end of that feature and exposed the other. Both bundled Bots refuse to start without OPENAI_API_KEY, so somebody who filled in an address for an Ollama or a vLLM got two dead containers complaining about a key their own server does not have. The row's summary names Ollama and vLLM by name; they were the two cases it would not serve. A base URL is a model, and its key belongs to it. Set, it means any endpoint speaking that API, which is what the variable's own comment has always said, so the startup check now asks for a key only when nothing else was named. Plain OpenAI still refuses without one, which is the case the check was written for, and the other two providers have no base URL to be named by so neither changes. The SDK insists on a string even when the endpoint ignores it, so a named endpoint with no key is handed a placeholder rather than a client that cannot be constructed. The decision is a module in each Bot rather than a condition at module scope, because index.ts serves as it loads and a test cannot import it without binding a port. Same reason model-options.ts exists. Also the model name now reaches the bundled Bot. docker-compose.yml reads AGENT_BOT_MODEL for agent-bot, not BOT_MODEL, so that a model chosen for the framework Bot cannot silently take its tools away: that Bot writes /v1/chat/completions by hand and gpt-5.6-* rejects function tools there. The reasoning is about OpenAI's own catalogue and does not survive a custom endpoint, where the pin asked the person's own server for a gpt-5.5 it has never heard of. The name they typed is written to both, and cleared from both when they answer with something that names no model. * Send a placeholder key to an endpoint that reads no key The Bots no longer demand a key when a base URL names the endpoint, but a Bot image published before they learned that still does, and a deployment pulls the image the release pinned. So the keyless half of the compatible row would have stayed broken until the next release, on every machine. The OpenAI SDK every Bot is built on refuses to construct a client without a string, which is the whole reason a blank key kills them. Ollama, vLLM, LM Studio and llama.cpp all ignore the value, so a placeholder is sent instead of nothing and the endpoint that does not read it is none the wiser. Never treated as a credential: it is written in plain sight rather than put in the machine's store, because it is not one. A key somebody actually typed is used unchanged. * Serve the installed app without a development server "Show OpenBot" did nothing on a machine where the stack was up. The window said OpenBot is running, the button was there, and clicking it had no effect at all. Two faults, one behind the other. The app host process was dead. It is started through the package's `serve` script, which ran `vite preview` through `bun --bun` so that a machine with bun and no Node could start it: `node_modules/.bin/vite` begins with a Node shebang. But Vite's proxy calls `socket.destroySoon()` when an upstream response ends, and bun's sockets do not implement it, so the process died with a TypeError on the FIRST call the app made. It served its page, exited, and nothing was listening on 3010 from then on. The shell went on reporting a stack that was up, because the containers were. So the app is served by a small server of its own now. It serves a directory and forwards one prefix, which is all an install needs; a development server was never the right thing to be running in an installed application, as the shell's own comment about this process already said. No Node, no Vite at runtime, and the websocket upgrade the live screen needs is forwarded rather than answered with HTML. A miss under /assets is still a 404 rather than the page, because handing a script tag some HTML fails in the console instead of the network panel. Paths are normalised and confined to the directory: the deployment's .env sits two levels above it. And the button now shows what it was told. `show_openbot` already answered with "OpenBot is not answering on port 3010 yet, so there is nothing to show", and the click handler dropped it with `.catch(() => undefined)`. A true sentence was available and the window threw it away, which is why this looked like a dead button rather than a dead process. The Ask screen's copy of the same call always showed it. * Show the CopilotKit sign-in address, the way the plan sign-ins do Both plan sign-ins keep the URL they were given and put it on screen, with a comment saying why: an open that silently does nothing, or a machine with no registered browser, leaves somebody watching a spinner with no idea where they are meant to go. The CopilotKit sign-in discarded it, so that case had no way out at all. Found while driving setup on a machine whose browser is not the one in front of the person. * Stop the host processes on Windows, where Stop was leaving them running Stop took the containers down, reported success, and left OpenBot serving. Measured on Windows Server 2022: after it, the server still answered on 3001, the worker was still up, and both halves of the app still answered 200 on 3010. Only the five containers had gone. The handles this window holds cover only what this window started, and they are gone the moment it restarts, so a window stopping a stack an earlier one started holds nothing. That is the case `stop_processes_under` exists for, and its Windows arm returned 0 with a comment saying the host processes end with the session. They do not. They are found by the ports the deployment publishes now, which the shell already owns and already checks for clashes, and each is ended with its children: `bun run serve` starts the real server as a grandchild, so ending the process holding the port would leave that one behind. Only the app's and the server's ports, because the containers are Compose's to stop and killing whatever holds a published container port reaches into the engine's own plumbing. Also, an unreadable app manifest is no longer reported as an old deployment. A byte-order mark in front of package.json made serde_json refuse it, and the refusal was rendered as "the deployment is older than this version of OpenBot", which sends somebody looking for a newer installer over three bytes. Windows tooling writes that mark freely: Set-Content -Encoding UTF8 does. It is skipped, and a manifest that genuinely will not parse says so. * Find the host processes by pid file, so the worker is stopped too The port sweep freed 3001 and 3010 and left the worker running. It listens on nothing, so a sweep cannot see it, and its command line is identical to the server's: both are `bun --env-file=../.env src/index.ts`, differing only by working directory, which Windows will not tell you cheaply. So the pids are written beside the logs when the processes start, and Stop reads them. That is also the honest fix for the case the sweep was standing in for: the handles a window holds die with the window, and everything else about a running stack survives it, so a restarted window Stopping a stack an earlier one started had nothing to work with. Now it has. The sweep stays as a second pass for a stack whose pid file is gone. The parse is its own function with a test on real netstat output, because reading five columns as four is what made the first attempt report success while leaving everything running: the foreign address was taken for the state and the state for the pid, so nothing ever matched. * Ask the credential store once per run, not once per screen Four Keychain dialogs, every time the setup screen mounted, each needing a click before the window would go on. Navigating between setup and OpenBot asked four more times. macOS authorizes every individual read of a stored password unless the application is signed with an identity the item's ACL already trusts. A development build is re-signed on every compile, so its ACL never matches and every read is a prompt; the wizard reads four secrets to arrive filled in, and it reads them on mount. The store is now asked once per name per process and the answer is held in memory. Absence is cached too, or a machine with no stored credential is asked on every mount for something that was never there. Writes go through the cache and forgetting clears it, so the two cannot disagree. This does not remove the prompts on a first run, and nothing in this process can: the decision belongs to the operating system and to the signature. A signed and notarised build is granted once and never asked again, which is the real fix and belongs to the release. * Scope the Keychain's service name to the Keychain CI builds on Linux with `-D warnings`, and there `SERVICE` is dead: only macOS has a service name to file a password under. Windows keys its DPAPI blobs by filename and the Linux fallback is a file in the config directory, so both ignored it. macOS never noticed, because there it is used three times. * Fix credential migration purge ordering Persist migrated secrets to the vault before rewriting .env with those secret keys purged. This keeps old file copies durable if any nonempty vault write fails, while preserving the successful purge path and empty-secret forgetting behavior. * Fix picked harness run routes Register Agno at /agui and LlamaIndex at /run when generating PICKED_HARNESS_URL, while keeping health_path as readiness metadata. Red: cargo test --lib harness::tests and env::tests failed because PickedHarness had no run_path field. Green: cargo test --lib harness::tests; cargo test --lib env::tests; bun test tests/compose.test.ts. * Verify Windows host identity before cleanup Record Windows host process identity at startup and require the live PID, executable, command line, and creation date to match before selecting taskkill roots. Restrict the netstat listener sweep to verified OpenBot process trees so unrelated listeners and stale or reused recorded PIDs are skipped while legitimate recorded children remain eligible for cleanup. * fix: load auth headers for remote mastra agents * fix(desktop): fetch deployment before harness image resolution * Fix CrewAI Bot role preservation * fix server mastra middleware cloning * fix app proxy websocket session headers * Normalize blank OpenAI base URL in Python harness * fix: preserve openbot context for mastra runs * fix: keep failed channel history reads neutral * fix anthropic picked harness provider config * fix: mount chatgpt token store directory * Fix provider-scoped plan sign-ins * Fix Ask model recovery state * test: isolate plugin-store audit refusal row * test: update remote agent wrapper assertions * fix: use native Mastra transport for desktop ask * test: isolate keychain regressions from real stores * fix desktop empty project retry * fix: split passive and interactive vault reads * Fix Mastra receiver OpenBot instructions * Add Python harness CI regressions * test: cover saved plan restoration * fix(desktop): carry byo agent endpoint through setup * style(desktop): format saved credential notice * test(desktop): isolate Rust temp roots * test(server): stub vertex provider in bun preload * test: satisfy vault cache clippy slice refs * fix desktop shutdown root retention Capture the active shell root before clearing it during shutdown so cleanup keeps targeting the running deployment root instead of the fallback/default root. Call-site audit: stop_stack still enters stop_everything with the caller fallback; tray/menu Stop still passes default_root as fallback; RunEvent::Exit now uses default_root only as fallback. Both stop_everything and RunEvent::Exit call shutdown_root before stop_processes_under and stack::down, so the selected non-default root reaches the external cleanup boundary while shell.root is cleared promptly to disable supervision. Windows cleanup identity checks in stack.rs were not changed. * fix: isolate compatible endpoint api keys * Reject missing saved model API keys * Add saved key startup boundary fixture * fix(desktop): scope saved setup state to root edits * Fix compatible endpoint scheme validation * Repair engine readiness compose install * fix: preserve ask body read errors * fix(desktop): read managed bot logs for fallback ask failures * fix(desktop): validate compatible endpoint credentials * fix(desktop): invalidate configured loads on root edits * fix(langgraph): preserve providers for colon-bearing model IDs * fix(desktop): omit API keys from plan choices * fix(desktop): detect IPv6-only port conflicts * fix(desktop): reject invalid provider endpoint hosts * fix(desktop): require a usable BYO agent URL host * fix(desktop): surface Windows setup detection failures * fix(desktop): protect Linux vault files before writing secrets * fix(app): contain static paths within the dist directory * fix(app): refuse malformed static URL encodings * fix(desktop): block disabled Virtual Machine Platform * fix(app): wait for upstream websocket handshake * fix(desktop): give administrators actionable WSL setup instructions * fix(desktop): recognize localized WSL kernel versions * fix(app): refresh channel history notices after Bot activity * fix(desktop): keep passive credential hydration out of protected storage * fix(server): report unavailable connected-vendor guidance * test(crewai): restore role test authentication token after teardown * fix(desktop): use the OpenBot orb for app and tray icons * fix(desktop): reject incomplete DPAPI stdin writes * fix(server): report standing instruction read failures * fix(desktop): restore hidden window on macOS reopen * fix(desktop): preserve unreadable settings files * fix(desktop): restore second instances outside the async listener * fix(server): report skill selection read failures * fix(mastra): default blank model identifiers * fix(desktop): report Linux credential read failures * fix(agent-mastra): validate mastra listen port * fix(server): diagnose discovery audit write failures * fix: normalize crewai blank model config * fix(desktop): fail closed on compose service inspection errors * fix(desktop): harden compose service inspection proofs * fix: validate configured ChatGPT auth file * test(server): restore tool selection model env * test(server): prove tool selection env restoration * test(server): harden tool selection env proof * fix(agent-crewai): project provider chat messages * fix: route google provider through google genai * fix(server): reject spoofed mastra governance context * fix(agent-crewai): forward caller tools to litellm * fix: pass anthropic base url to picked harness * test(server): clean up custom server upsert credential * fix(server): restrict tombstones to live channels * fix: surface process cleanup failures * fix: propagate cleanup command failures * test: verify every secret-bearing compose port binding * test(server): prove the current tool refusal audit payload * fix: reject unavailable Windows cleanup evidence * fix: retain Windows pidfile after incomplete cleanup * fix: persist host pidfiles with strict atomic replacement * test(server): restore borrowed Notion client credentials * fix(desktop): verify Unix process ownership before cleanup * fix(server): preserve deleted Bot refusal through runtime clones * fix(app): validate built serving ports with shared contract * fix(app): retain failed mount history notice * fix(app): restore durable history around shared message anchors * fix(desktop): suppress credential authorization UI during normal actions * fix(desktop): preserve existing installation encryption keys * fix(langgraph): default blank providers before model routing * fix(app): restore readable channel history on mount * test(server): bind plugin audit assertions to each invocation * fix(desktop): recover one refused Keychain operation * fix(desktop): restore noninteractive Windows vault invocation * fix(langgraph): select defaults for the configured provider * fix(desktop): retire failed starts before retry cleanup * Use root-scoped file credential storage * Add credential root boundary proof tests * test(server): prove fixture listener teardown * fix(server): preserve handoff initiator through delivery * test(server): bind handoff delivery signer proof * fix(desktop): retain selected root for menu stop * fix(langgraph): bind and execute forwarded tools * fix app bot default to picked harness * test: tighten bot default route fixture * test: document bot route wiring cast * fix(app): ignore self echoed channel activity * fix(app): treat malformed thread history as unavailable * fix(server): filter reserved mastra context fields * fix(app): reject decoded nul static paths * fix(server): preload EventSource for production loader * test(server): accept windows loader db boundary * fix(scripts): stop legacy server on restart * fix(langgraph): preserve parallel tool call stream lifecycles * fix(app): bound stalled thread history reads * style(server): wrap loader boundary assertion * fix landing default agent selection * fix(computer): skip unavailable optional SPIRE sockets * fix(desktop): validate BYO harness URLs before persistence * fix desktop start on dead compose services * fix(desktop): select bundled bot services by provider * fix(desktop): report unreadable DPAPI credential files * fix(desktop): distinguish DPAPI read diagnostics * fix(desktop): clear saved model intent for no model * fix(desktop): preserve private storage file boundaries * fix(desktop): detect localized WSL1 status * fix(desktop): parse WSL default version label precisely * fix(desktop): read WSL default version from registry * fix(desktop): fail loud on WSL registry read errors * fix(desktop): keep WSL registry probe strict and readable * docs(desktop): clarify WSL registry version probe * fix(desktop): fail loud on Windows MSI cleanup errors * test(desktop): gate MSI process proof to Unix * fix(desktop): keep Podman command failure evidence * test(desktop): gate Podman process fixture import * fix(desktop): reject line breaks in dotenv settings * fix(compose): map selected harness callback host * fix(desktop): fail closed on partial Windows host ownership * fix(desktop): retire host children after ownership record failure * fix(desktop): retain host ownership when cleanup fails * fix(desktop): verify root ownership before adopting running stack * fix(desktop): require Unix server port ownership before adoption * fix(desktop): restore endpoint-scoped saved model credentials * fix(app): report bot roster load failures * fix bot route hidden agent detail lookup * fix bot route detail cache collision * fix channel new agent load errors * fix channel new stale detail error state * fix(server): cancel pending agent builds on Stop * fix home fallback agent selection * fix same-timestamp activity refresh * fix LangGraph test imports from repository root * fix empty history refresh retry * fix no-text run activity reporting * fix cached channel activity restore * fix desktop Start for saved compatible endpoints * fix empty bot agent query default * fix BYO startup to skip the local harness service * fix bundled agent registry eligibility across provider changes * Support container endpoint override for compatible models * fix(app): encode agent IDs in API paths * fix(server): authenticate picked harness without bundled bot * fix(server): abort wrapped remote agent runs * fix(server): canonicalize managed endpoint identity * test(compose): verify compatible endpoint resolution * fix(desktop): retain ownership after partial host launch * fix(server): scope tenant package grant cleanup * fix(server): protect tenant package channel ownership * fix(desktop): scope computer shutdown to deployment namespace * fix(server): use selected compatible model for built-in agents * test(supervisor): isolate health-capable Docker lifecycle fixtures * fix(server): honor explicit openai compatible model * fix(desktop): restore only the selected owned deployment * fix(desktop): verify app listener ownership before adoption * fix(desktop): keep Quit responsive during cleanup * fix(desktop): cancel initial startup when Stop retires its run * fix(desktop): verify Windows parent process creation order * test(desktop): make native fixtures portable and isolate their environment * fix(desktop): confirm Unix descendants exit before stopping ancestors * fix(desktop): stop Windows process trees without stale port sweeps * fix(desktop): retry frozen dependency installs after partial failures * fix(desktop): stop supervisor before snapshotting computers * fix(desktop): parse every Compose published port row * fix(desktop): require current API readiness before startup succeeds * style(desktop): remove blank-comment trailing whitespace Remove only eight spaces from the blank JSX comment line in HarnessPicker. The blank line and all non-whitespace source are unchanged. Validation: Biome 2.5.10 scoped format check; git diff --check against origin/main; git diff -w HEAD empty before commit. No symbols or call sites change; no behavior tests required. * test(desktop): preserve PowerShell in restricted Windows fixtures * test(desktop): serve readiness fixtures on both loopbacks * fix(desktop): resolve setup navigation from Tauri configuration * test(desktop): name native fixture crates explicitly * test(desktop): use platform local origin for mock IPC * ci(desktop): run library and shell regressions on every platform * fix(desktop): include IPv6 listeners in Windows port ownership * fix(desktop): dispatch Stop shutdown on a blocking worker * fix(desktop): retain unresolved legacy Windows PID evidence * fix(desktop): preserve supervision through failed Start preflight * fix(desktop): clean up owned listener test fixture Return the compiled listener directory to its single fixture caller and remove it after successfully reaping the listener. Preserve the existing real-process ownership regression. * fix(desktop): record Windows supervisor replacements before success Persist a replacement's complete direct-child identity while its live Child remains held, before the supervisor can announce restart success. Preserve prior role records without re-recording dead or reaped children, allowing server and app to restart sequentially after both have exited. Retain handles and durable evidence on every recording failure. Callsite audit: restart_host_process_with is the only production caller of replace_windows_host_process_with. Its existing children lock still covers generation/root validation, spawn, recording and publication, with the same post-publication generation check used by concurrent Stop. The supervisor reports started again only after Ok(true). Unix replace_host_process, initial Windows recording, durable-write semantics, held-child cleanup supplementation, and tray/second-instance ownership probes remain unchanged. Validation: three ordinary replacement regressions pass, including stale PID reuse, other-role retention, sequential dual death and recording refusal. Source-bound Windows-arm restart-to-ownership proof fails three expected baseline cases and passes all four final controls. The failure proof reaches the actual Stop handle cleanup. Twenty-nine existing Windows ownership and cleanup regressions pass. cargo fmt --check and strict Clippy all-targets pass. Native real-listener proof is supplied separately for root-owned validation. Finding: D89-MAIN1-001-windows-restart-records * fix(desktop): retain container ownership across failed Start retries * fix(desktop): attribute empty answers to the selected endpoint * fix(desktop): quiesce Unix launchers before descendant cleanup * fix(desktop): retain runtime affinity for container cleanup * test(desktop): cover unavailable-engine shutdown recovery * fix(desktop): retain recovery until the owned run is restored * test: preserve CI test stderr before exit * test: keep CI stderr regression deterministic * test: fix desktop ci fixture portability * test: run ChatGPT token writer as host user * test: fix inline podman runtime fixture * test: bind plugin audit JSON filters * test: stop channel new mock leaks * ci: install desktop deps before root tests * fix: retain menu stop failures through setup navigation * test: repair merged plugin OAuth fixtures * fix: preserve ChatGPT token store ownership on refresh * fix: retain quit cleanup notice for setup recovery * test(desktop): model local Podman fixture affinity
1 parent 8819f18 commit c21ba2b

201 files changed

Lines changed: 47837 additions & 1962 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.env.example‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -160,6 +160,11 @@ OPENAI_API_KEY=
160160
#
161161
# OPENAI_BASE_URL=
162162

163+
# Optional. Leave empty unless containers need a different route to the same compatible endpoint,
164+
# such as a locally hosted model whose host URL is localhost but whose Compose-network URL is a
165+
# service name. Empty means containers use OPENAI_BASE_URL too.
166+
# OPENAI_CONTAINER_BASE_URL=
167+
163168
# The same for the other two providers, under the names the API server already reads. They are
164169
# different APIs rather than different URLs for this one, so each has its own.
165170
# ANTHROPIC_BASE_URL=

‎.github/published-images.json‎

Lines changed: 19 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1 +1,19 @@
1-
["agent-computer", "supervisor", "agent-bot", "agent-langgraph", "server"]
1+
[
2+
"agent-computer",
3+
"supervisor",
4+
"agent-bot",
5+
"agent-langgraph",
6+
"server",
7+
"agent-crewai",
8+
"agent-llamaindex",
9+
"agent-agno",
10+
"agent-langgraph-agui",
11+
"agent-adk",
12+
"agent-pydantic-ai",
13+
"agent-microsoft",
14+
"agent-claude-sdk",
15+
"agent-strands",
16+
"agent-ag2",
17+
"agent-langroid",
18+
"agent-mastra"
19+
]

‎.github/workflows/ci.yml‎

Lines changed: 42 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -215,6 +215,14 @@ jobs:
215215
# Bot's own tree and not in the root one.
216216
- run: bun install --frozen-lockfile
217217
working-directory: agent-langgraph
218+
# Same for the Mastra Bot: root test discovery imports its receiver tests, which load Mastra
219+
# from that Bot's own dependency tree.
220+
- run: bun install
221+
working-directory: agent-mastra
222+
# Same for the desktop app: root test discovery imports its React tests, whose JSX runtime is
223+
# pinned by the desktop lockfile rather than the root one.
224+
- run: bun install --frozen-lockfile
225+
working-directory: desktop
218226
# Not the db:migrate script: that one loads ../.env, which does not exist in CI. DATABASE_URL
219227
# comes from the job env instead, which drizzle.config.ts already reads.
220228
- run: bunx drizzle-kit migrate --config=drizzle.config.ts
@@ -223,6 +231,39 @@ jobs:
223231
# files before their tests are registered.
224232
- run: bun run test:ci
225233

234+
python-harness:
235+
name: python harness regressions
236+
runs-on: ubuntu-latest
237+
steps:
238+
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
239+
with:
240+
persist-credentials: false
241+
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
242+
with:
243+
python-version: "3.12"
244+
- uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2.2.0
245+
with:
246+
bun-version: 1.3.14
247+
- name: CrewAI provider-boundary regression
248+
run: |
249+
set -euo pipefail
250+
python -m venv .venv-crewai
251+
. .venv-crewai/bin/activate
252+
python -m pip install --requirement agent-crewai/requirements.txt --requirement agent-crewai/requirements-test.txt
253+
python -m pytest agent-crewai/tests -q
254+
- name: LangGraph AG-UI provider-boundary regressions
255+
run: |
256+
set -euo pipefail
257+
python -m venv .venv-langgraph-agui
258+
. .venv-langgraph-agui/bin/activate
259+
python -m pip install --requirement agent-langgraph-agui/requirements.txt --requirement agent-langgraph-agui/requirements-test.txt
260+
python -m pytest agent-langgraph-agui/tests -q
261+
- run: bun install --frozen-lockfile
262+
- run: bun test tests/compose.test.ts
263+
- run: docker compose --env-file /dev/null --profile harness config --format json >/dev/null
264+
env:
265+
PICKED_HARNESS_IMAGE: openbot-agent-langgraph-agui:test
266+
226267
build:
227268
name: build
228269
runs-on: ubuntu-latest
@@ -414,7 +455,7 @@ jobs:
414455
name: verify
415456
runs-on: ubuntu-latest
416457
if: always()
417-
needs: [static, deployables, chart, test, build, migrations, image, component-dockerfiles]
458+
needs: [static, deployables, chart, test, python-harness, build, migrations, image, component-dockerfiles]
418459
steps:
419460
- name: Require every check
420461
env:

‎.github/workflows/desktop.yml‎

Lines changed: 7 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,7 @@ name: Desktop
22

33
# The shell is the one thing here that cannot be proved by running it on Linux: it is a macOS app, a
44
# Windows app and a Linux app built from one tree, and the ways they differ are exactly the ways
5-
# this fails. So it builds on all three, on every change to it.
5+
# this fails. So it tests and builds on all three, on every change to it.
66
on:
77
pull_request:
88
paths: ["desktop/**", ".github/workflows/desktop.yml"]
@@ -19,9 +19,7 @@ concurrency:
1919
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
2020

2121
jobs:
22-
# Fast, and the only job that runs the assertions about what the shell writes into `.env` and
23-
# which socket it names. Those are the parts that carry what the platforms taught us, and they are
24-
# plain Rust: no window, no engine, no waiting.
22+
# Fast formatting and lint checks. Rust regression tests run once per platform below.
2523
core:
2624
name: core
2725
runs-on: ubuntu-latest
@@ -44,8 +42,6 @@ jobs:
4442
working-directory: desktop/src-tauri
4543
- run: cargo clippy --all-targets -- -D warnings
4644
working-directory: desktop/src-tauri
47-
- run: cargo test --lib
48-
working-directory: desktop/src-tauri
4945

5046
# The three artifacts. Not `bundle`, which signs and notarizes: that is S7 and needs certificates
5147
# this workflow deliberately does not hold. This proves the tree builds into an app on each
@@ -95,6 +91,11 @@ jobs:
9591
# which a compile would miss.
9692
- run: bun run tauri build
9793
working-directory: desktop
94+
# Frontend assets and platform dependencies are ready after the packaged build.
95+
# Include main.rs regressions as well as lib.rs; ignored live tests remain opt-in.
96+
- name: Rust regression tests
97+
run: cargo test --locked --lib --bins
98+
working-directory: desktop/src-tauri
9899
# Keep what was built. Without this the only way to try an installer is to build one on
99100
# the machine you are trying it on, which is not what anybody installs.
100101
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1

‎agent-adk/Dockerfile‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
# Google ADK, as a Bot. Python rather than Bun because that is what the integration is published in,
2+
# and the protocol is the only thing a Bot has to share with the others.
3+
FROM python:3.12-slim
4+
5+
WORKDIR /app
6+
7+
# Dependencies before source, so editing the Bot does not re-resolve the framework.
8+
COPY agent-adk/requirements.txt ./
9+
RUN pip install --no-cache-dir -r requirements.txt
10+
11+
COPY agent-adk/src ./src
12+
13+
ENV PORT=4208
14+
EXPOSE 4208
15+
# 0.0.0.0, not `::`. uvicorn binds `::` as IPv6-only, with no v4-mapped addresses, so a container
16+
# started that way refuses 127.0.0.1: the Compose healthcheck never passes and the server cannot
17+
# reach the Bot. Measured, not assumed. Inside a container this is the container's own namespace,
18+
# and Compose is what decides which host addresses it is published on.
19+
CMD ["uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "4208"]

‎agent-adk/requirements.txt‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,6 @@
1+
ag-ui-adk
2+
google-adk
3+
litellm
4+
fastapi
5+
python-multipart
6+
uvicorn[standard]

‎agent-adk/src/main.py‎

Lines changed: 54 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,54 @@
1+
"""Google ADK as a Bot, through `ag_ui_adk`, which AG-UI maintains.
2+
3+
ADK is Gemini-first and model-agnostic after that, so the provider stays the person's choice: ADK
4+
reads LiteLLM model strings, and OpenBot writes the one it was told.
5+
"""
6+
7+
import os
8+
9+
from ag_ui_adk import ADKAgent, add_adk_fastapi_endpoint
10+
from fastapi import FastAPI, Request
11+
from fastapi.responses import JSONResponse
12+
from google.adk.agents import Agent
13+
from google.adk.models.lite_llm import LiteLlm
14+
15+
TOKEN_HEADER = "x-openbot-agent-token"
16+
17+
18+
def _model_id() -> str:
19+
provider = (os.environ.get("BOT_PROVIDER") or "openai").strip()
20+
model = (os.environ.get("BOT_MODEL") or "gpt-4o-mini").strip()
21+
return model if "/" in model else f"{provider}/{model}"
22+
23+
24+
app = FastAPI()
25+
26+
27+
@app.middleware("http")
28+
async def refuse_without_the_server_token(request: Request, call_next):
29+
if request.url.path != "/health":
30+
expected = (os.environ.get("MANAGED_AGENT_TOKEN") or "").strip()
31+
offered = (request.headers.get(TOKEN_HEADER) or "").strip()
32+
if not expected or offered != expected:
33+
return JSONResponse({"error": "unauthorised"}, status_code=401)
34+
return await call_next(request)
35+
36+
37+
@app.get("/health")
38+
async def health():
39+
return {"ok": True, "harness": "google-adk"}
40+
41+
42+
add_adk_fastapi_endpoint(
43+
app,
44+
ADKAgent(
45+
adk_agent=Agent(
46+
name="openbot",
47+
model=LiteLlm(model=_model_id()),
48+
instruction="Answer the question you are asked, briefly and correctly.",
49+
),
50+
app_name="openbot",
51+
user_id="openbot",
52+
),
53+
path="/",
54+
)

‎agent-ag2/Dockerfile‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,19 @@
1+
# AG2, as a Bot. Python rather than Bun because that is what the integration is published in,
2+
# and the protocol is the only thing a Bot has to share with the others.
3+
FROM python:3.12-slim
4+
5+
WORKDIR /app
6+
7+
# Dependencies before source, so editing the Bot does not re-resolve the framework.
8+
COPY agent-ag2/requirements.txt ./
9+
RUN pip install --no-cache-dir -r requirements.txt
10+
11+
COPY agent-ag2/src ./src
12+
13+
ENV PORT=4210
14+
EXPOSE 4210
15+
# 0.0.0.0, not `::`. uvicorn binds `::` as IPv6-only, with no v4-mapped addresses, so a container
16+
# started that way refuses 127.0.0.1: the Compose healthcheck never passes and the server cannot
17+
# reach the Bot. Measured, not assumed. Inside a container this is the container's own namespace,
18+
# and Compose is what decides which host addresses it is published on.
19+
CMD ["uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "4210"]

‎agent-ag2/requirements.txt‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,4 @@
1+
ag2[ag-ui,openai]
2+
fastapi
3+
python-multipart
4+
uvicorn[standard]

‎agent-ag2/src/main.py‎

Lines changed: 39 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,39 @@
1+
"""AG2 as a Bot. AG-UI is an extra in AG2's own package, so `ag2[ag-ui]` is the dependency."""
2+
3+
import os
4+
5+
from ag2 import Agent
6+
from ag2.ag_ui import AGUIStream
7+
from ag2.config import OpenAIConfig
8+
from fastapi import FastAPI, Request
9+
from fastapi.responses import JSONResponse
10+
11+
TOKEN_HEADER = "x-openbot-agent-token"
12+
13+
agent = Agent(
14+
name="openbot",
15+
prompt="Answer the question you are asked, briefly and correctly.",
16+
config=OpenAIConfig(model=(os.environ.get("BOT_MODEL") or "gpt-4o-mini").strip()),
17+
)
18+
stream = AGUIStream(agent)
19+
20+
app = FastAPI()
21+
22+
23+
@app.middleware("http")
24+
async def refuse_without_the_server_token(request: Request, call_next):
25+
if request.url.path != "/health":
26+
expected = (os.environ.get("MANAGED_AGENT_TOKEN") or "").strip()
27+
offered = (request.headers.get(TOKEN_HEADER) or "").strip()
28+
if not expected or offered != expected:
29+
return JSONResponse({"error": "unauthorised"}, status_code=401)
30+
return await call_next(request)
31+
32+
33+
@app.get("/health")
34+
async def health():
35+
return {"ok": True, "harness": "ag2"}
36+
37+
38+
# `build_asgi` rather than a helper: AG2 hands back a plain ASGI app to mount.
39+
app.mount("/", stream.build_asgi())

0 commit comments

Comments
 (0)