OmniRoute

AI

Never stop coding. Free MIT AI gateway: one endpoint, 330+ providers (90+ free), 1200+ models — Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 320+ contributors

Latest v3.8.49 · by diegosouzapwWebsitediegosouzapw/OmniRoute

Release activity

Release activity — 10 releases across 8 days since Jun 29, 2026. Each cell is one day; darker means more releases that day. Nothing is recorded before Jun 29, 2026. Older weeks are hidden at this screen width.
MayJunJulAug
SundayNo releases on Jul 5, 2026No releases on Jul 12, 2026No releases on Jul 19, 2026No releases on Jul 26, 2026No releases on Aug 2, 2026No releases on Aug 9, 2026
Monday2 releases on Jun 29, 20261 release on Jul 6, 20262 releases on Jul 13, 2026No releases on Jul 20, 2026No releases on Jul 27, 2026No releases on Aug 3, 2026No releases on Aug 10, 2026
Tuesday1 release on Jun 30, 20261 release on Jul 7, 2026No releases on Jul 14, 2026No releases on Jul 21, 2026No releases on Jul 28, 2026No releases on Aug 4, 2026No releases on Aug 11, 2026
WednesdayNo releases on Jul 1, 2026No releases on Jul 8, 2026No releases on Jul 15, 2026No releases on Jul 22, 2026No releases on Jul 29, 2026No releases on Aug 5, 2026No releases on Aug 12, 2026
Thursday1 release on Jul 2, 2026No releases on Jul 9, 2026No releases on Jul 16, 2026No releases on Jul 23, 20261 release on Jul 30, 2026No releases on Aug 6, 2026No releases on Aug 13, 2026
FridayNo releases on Jul 3, 2026No releases on Jul 10, 2026No releases on Jul 17, 2026No releases on Jul 24, 2026No releases on Jul 31, 2026No releases on Aug 7, 2026
Saturday1 release on Jul 4, 2026No releases on Jul 11, 2026No releases on Jul 18, 2026No releases on Jul 25, 2026No releases on Aug 1, 2026No releases on Aug 8, 2026

10 releases since Jun 29, 2026, busiest day 2

Changelog

v3.8.49

Added 20
  • Generalize ensureThinkingBudget to all providers
  • Add effort-tier aliases for glm-5.2 and mimo-v2.5
  • Add curated OpenRouter embeddings catalog with specialty merge
  • Add opt-in auto-ping to keep Codex quota windows warm
  • Add Agnes AI native provider support
  • Allow disabling `:` comment heartbeats via SSE configuration

All 1383 entries from this cycle are listed below, one line each — descriptions are trimmed to fit GitHub's 125,000-character release body. Full wording, context and links: CHANGELOG.md. Living section — regenerated 2026-07-19 from all 306 cycle commits (bump 2c62333b0 → tip). Bullets carry the merged PR and its author; direct pushes listed separately. Finalized at the v3.8.49 release.

✨ New Features
  • feat: generalize ensureThinkingBudget to all providers +… (#6979) — @rafaumeu

  • feat(6922): effort-tier aliases for glm-5.2 & mimo-v2.5 on… (#6987) — @rafaumeu

  • feat(providers): curated OpenRouter embeddings catalog + specialty merge… (#6994)

  • feat(quota): opt-in auto-ping to keep Codex quota windows warm (#6995)

  • feat(providers): add Agnes AI native provider support (#7035) — @HouMinXi

  • feat(sse): allow disabling : comment heartbeats via… (#7036) — @xier2012

  • feat(perf): add performance.mark/measure to SSE pipeline +… (#7045) — @oyi77

  • feat(providers): add Dahl free inference provider (#7062) — @growab

  • feat(ci): boot-smoke the packed npm tarball (check:pack-boot,… (#7086)

  • feat(ci): hotfix fast-lane + tests-only E2E skip (WS3.1) (#7088)

  • feat(ci): continuous release-green — on-push quick gate + 3x/day… (#7089)

  • feat(ci): duration-balanced E2E shards via LPT bin-packing (WS4.1) (#7090)

  • feat(ci): TypeScript 7 native shadow for typecheck:core (WS4.2,… (#7091)

  • feat(release): npm staged publishing + pre-publish boot-smoke (WS1.3) (#7092)

  • feat(release): post-publish verifier — clean-container install + boot… (#7109)

  • feat(ci): Mergify merge queue + manual-train fallback runbook… (#7112)

  • feat(ci): Windows leg for Electron prepare smoke (WS1.5) (#7113)

  • feat(ci): Codecov patch coverage (informational) + fix missing… (#7114)

  • feat(sidecar): support conditional provider manifest refresh (#7130) — @KooshaPari

  • feat(homolog): real-environment E2E homologation suite (npm run… (#7133)

  • feat(usage): add Codex reset credit picker (#7154) — @JxnLexn

  • feat(ci): Trunk Flaky Tests uploads for vitest + Playwright E2E… (#7175)

  • feat(ci): Trunk Flaky Tests upload on the fast-path vitest job… (#7205)

  • feat(kiro): register GPT-5.6 Sol/Terra/Luna model family (#7209)

  • feat(dashboard): show Codex plan label in provider and quota views (#7210)

  • feat(dashboard): add reorder connections by availability button (#7211)

  • feat(dashboard): add 180D and 365D usage/cost analytics periods (#7213)

  • feat(api): add Vary: Accept-Encoding to token-authenticated /v1*… (#7217)

  • feat(api): expose GET /api/usage/model-latency-stats (#7218)

  • feat(dashboard): add compression-mode selector to Context & Cache combos… (#7219)

  • feat(sse): route GitHub Copilot Claude models through native… (#7223)

  • feat(mitm): add Antigravity reasoning-effort overrides (#7228)

  • feat: replace free-text model inputs with hidePaid-aware… (#7229)

  • feat: editable ComfyUI base-URL field + per-connection… (#7232)

  • feat(sse): add optional-enum null-omission idiom for codex… (#7233)

  • feat(sse): preserve tools/tool_choice for tool-bearing requests… (#7235)

  • feat(api): accept x-goog-api-key header for client-facing auth (#7236)

  • feat(sse): add native xAI Grok Imagine video generation provider (#7238)

  • feat: add Type filter and easiest-first sort to Free Provider… (#7240)

  • feat(cli): add Grok Build CLI tool setup (~/.grok/config.toml) (#7241)

  • feat(provider): add Chenzk API OpenAI-compatible gateway (#7246)

  • feat(providers): let custom connections opt into prompt-cache capability (#7257)

  • feat(db): include xp_audit_log in automatic retention/prune (#7260)

  • feat(api): structured X-Routing-Fallback-Reason header for relay… (#7262)

  • feat(compression): support RTK TOML schema v1 filters (#7281) — @JxnLexn

  • feat: add principal-scoped CCR MCP lifecycle (#7282) — @JxnLexn

  • feat(morph): refresh curated models (#7314) — @backryun

  • feat(issue-agent): surface RecordedTriageTimeoutError as 504 (#7315) — @KooshaPari

  • feat(incident-response): structured incident response templates (#7334) — @KooshaPari

  • feat(providers): add xAI OAuth PKCE provider (#7399) — @fenix007

  • feat(models): advertise Claude reasoning-effort variants in /v1/models (#7497) — @thepigdestroyer

  • feat(kimi): sync Code, Web, and Moonshot providers (#7531) — @backryun

  • feat(resilience): guard OmniRoute peer routing loops (#7555) — @isiahw1

  • feat: add Mixedbread AI as embeddings provider (#7595)

  • feat(providers): add Rev AI speech-to-text provider (#7596)

  • feat: add Freepik (Magnific Mystic) image generation provider (#7597)

  • feat(sse): add DeepInfra as a video-generation provider (#7598)

  • feat(providers): add Felo chat-aggregator provider (#7599)

  • feat(sse): add Notion AI Web (Unofficial/Experimental) provider (#7600)

  • feat: add FreeTheAi as OpenAI-compatible gateway provider (#7602)

  • feat: add Gladia as an async speech-to-text provider (#7603)

  • feat: add EdgeTTS audio-tts provider (#7605)

  • feat(video): add Novita AI as video-generation provider (#7606)

  • feat: add Segmind image+video provider (#7608)

  • feat: add Microsoft Designer as image provider (#7609)

  • feat: per-model default reasoning_effort + no-think none on… (#7631)

  • feat(sse): per-model upstream header-response timeout override (#7632)

  • feat(dashboard): in-product guidance for prompt compression engines (#7634)

  • feat(usage): add TTFT/E2E-latency/tokens-per-second to model latency… (#7635)

  • feat: import providers from CSV/JSON file (#7636)

  • feat: confirm before removing a single connection (#7640)

  • feat(sse): honor excluded models in no-auth auto-combo candidate… (#7646)

  • feat(providers): add g4f.space no-key gateway… (#7647)

  • feat: rate-limit queue admission control (maxQueueDepth + 15s… (#7649)

  • feat(sse): generalize session affinity TTL to all providers (#7650)

  • feat: OpenRouter quota tracking (key/credits + free-window… (#7651)

  • feat(sse): quota tracking for AgentRouter, v0 (Vercel), FreeModel… (#7653)

  • feat(providers): Speechmatics STT, gTTS, VibeProxy preset (#6659, #6667,… (#7655)

  • feat(api): route Google AI Studio Imagen through… (#7656) — @danscMax

  • feat(auth): OIDC as optional dashboard admin login gate (password… (#6973) — @mikolaj92

  • feat(api): add pagination params to 8 DB modules + recharts… (#7046) — @oyi77

  • feat(proxy): operator-level proxy subscriptions (Karing-style) —… (#7299) — @xier2012

  • feat(grok-cli): align with official Grok Build client (#7358) — @backryun

  • feat(providers): Complete GHE Copilot OAuth provider implementation (#7546) — @hppsc1215

  • feat(guardrails): add CredentialMaskerGuardrail for API key/secret… (#7683) — @Securiteru

  • feat(perplexity): refresh provider integrations (#7687) — @backryun

  • feat(providers): notion-web live model discovery via getAvailableModels (#7696) — @artickc

  • feat(providers): add proactive cf_clearance/User-Agent hint to grok-web… (#7713)

  • feat: add live gRPC-web quota fetcher for grok-cli (#7714)

  • feat(api): add opt-in auto-sync scheduler for free-proxy sources (#7716)

  • feat(dashboard): show proxy name in badge, sort saved-proxy picker,… (#7720)

  • feat(cli): add auth export command for decrypted provider… (#7724)

  • feat(oauth): accept full ChatGPT session JSON for Codex manual import (#7725)

  • feat(sse): add nvidia NIM local RPM budget + concurrency cap (#7726)

  • feat(gemini-web): emulate OpenAI tool calling via the webTools prompt shim (#7727)

  • feat(services): introduce pluggable service-provider contract, migrate… (#7730)

  • feat(mitm): root-CA + per-host leaf certs for AgentBridge static… (#7731)

  • feat(providers): add hailuo-web (MiniMax web) chat provider (#7734)

  • feat: browser login for Grok Build provider (#7735)

  • feat(routing): wire interceptFetch tool interception into the chat… (#7736)

  • feat(sse): add X-OmniRoute-Decision routing trace header (#7765)

  • feat(providers): zai-web live model discovery with local-catalog fallback (#7766)

  • feat(api): sync upstream reasoning.supported_efforts into… (#7767)

  • feat(dashboard): pin Kimi providers first in category + official… (#7775)

  • feat(chaos+ponytail): parallel chaos-mode dispatch + ponytail output … (#7781) — @Moseyuh333

  • feat(perf): IC2 — cache provider connections by ID + lazy-decrypt… (#7787) — @oyi77

  • feat(quality): gate the free-tier headline so it can never silently… (#7798)

  • feat(providers): expose an explicit tier override for any provider… (#7838)

  • feat(routing): read-only auto/* candidate transparency + per-API-key… (#7839)

  • feat(catalog): map unmapped free tiers, add navy + aihorde, surface… (#7840)

  • feat(providers): add OpenRouter speech-to-text (audio transcription)… (#7861) — @Tasogarre

  • feat(qwen): add Qwen3.8 Max Preview catalogs [Part 2/3] (#7874) — @backryun

  • feat: support Bun bundled SQLite runtime (#7878) — @Arul-

  • feat(providers): add 5 free-tier providers (ainative, aion, sealion,… (#7887)

  • feat(vnc-session): persistent noVNC browser login for web-cookie providers (#7892) — @Capslockb

  • feat(sse): add PromptQL playground provider (unofficial) (#7911) — @artickc

  • feat(cline): align ClinePass catalog and request protocol (#7914) — @backryun

  • feat: narrow mcp:connect scope + per-key HTTP tool-scope… (#7967)

  • feat: provider tab account search + mirrored top pagination (#7968)

  • feat: canonical numeric helpers + tier-1 (analytics) migration (#7969)

  • feat(sse): add HyperAgent (hyperagent.com) unofficial web provider (#7994) — @artickc

  • feat: copilot-m365-web tone-selected model variants (#7997)

  • feat(media): Adobe Firefly image + video generation provider (#8006) — @artickc

  • feat(routing): add prompt-cache affinity (#8008) — @JxnLexn

  • feat(compression): select model-aware tokenizers (#8009) — @JxnLexn

  • feat(compression): add Responses tool-output engine (#8010) — @JxnLexn

  • feat(dashboard): Kimi sponsor banner, Kimi Coding preset, official… (#8039)

  • feat(dashboard): make Codex quota card windows reflect reality (#8054) — @insoln

  • feat(compression): teach the model the CCR retrieve protocol on first… (#8063)

  • feat(compression): per-model/endpoint compression exclusion filter (#8064)

  • feat(providers): add CLOVA Studio, InternLM and Ant Ling API-key… (#8077) — @alvaretto

  • feat(codex): support reference image edits (#8122) — @xiaoyaner0201

  • feat(providers): add weekly quota tracking for grok-web (#8127) — @apoapostolov

  • feat(providers): add Sarvam AI, Writer Palmyra and PLaMo API-key… (#8161) — @alvaretto

  • feat: native Fish Audio TTS provider on /v1/audio/speech (#8164)

  • feat: zh-CN terminology glossary + consistency gate +… (#8166)

  • feat(providers): add Typhoon (Thailand) and Inception Mercury diffusion… (#8170) — @alvaretto

  • feat(sse): restrict auto-combo no-auth pool to allowlist… (#8183)

  • feat(sre): add tcp-close-analyzer.py for debugging… (#8208) — @hartmark

  • feat(settings): configurable model catalog cache TTL (#8219) — @oyi77

  • feat(github-models): refresh catalog and compatibility (#8225) — @backryun

  • feat(github): refresh Copilot model catalog (#8226) — @backryun

  • feat: classify grok-web Cloudflare anti-bot blocks + gated… (#8241)

  • feat(providers): weekly quota for xAI OAuth (Grok) (xai-oauth / xao)… (#8471) — @allanvb

  • feat(sse): every completion response now carries an… — @chirag127

  • feat(ws): the live-dashboard WebSocket server now auto-starts… — @ianriizky

  • feat(sse): configurable per-model upstream… (#6354)

  • feat(dashboard): Replace free-text model inputs in the Routing (web… (#6540)

  • feat(compression): new omniglyph engine (context-as-image) — renders… (#6556 #6661)

  • feat(sse): rate-limit request queue admission control —… — @chirag127

  • feat(sandbox): the skill sandbox gained a container-provider… — @KooshaPari

  • feat(oauth): accept the full ChatGPT session JSON (not… (#6636)

  • feat(sse): add 5 no-key g4f.space gateway providers —… — @chirag127

  • feat(sse): add DeepInfra as a video-generation provider… (#6653)

  • feat(providers): add Freepik (Magnific Mystic) API-key… (#6654)

  • feat(providers): add Rev AI speech-to-text provider… (#6655)

  • feat(providers): add Segmind as an image + video generation provider… (#6656)

  • feat(providers): add Gladia as an async speech-to-text… (#6657)

  • feat(video): add Novita AI as a video-generation… (#6658)

  • feat(providers): add Speechmatics as an STT provider — async batch… (#6659)

  • feat(providers): add Mixedbread AI as an embeddings… (#6660)

  • feat(providers): add Felo (felo.ai) as a free, no-signup, no-API-key… (#6666)

  • feat(providers): add gTTS (Google Translate TTS) as a free, no-signup… (#6667)

  • feat(sse): add EdgeTTS (Microsoft Edge "Read Aloud") as a free,… (#6668)

  • feat(providers): add FreeTheAi as an OpenAI-compatible… (#6670)

  • feat(sse): add Microsoft Designer as an unofficial… (#6672)

  • feat(providers): add Hailuo Web (hailuo-web) — a free, _token-based… (#6673)

  • feat(cli): new omniroute auth export command dumps DECRYPTED… (#6683)

  • feat(mitm): the AgentBridge static MITM server (server.cjs) can… (#6684)

  • fix(sse): skip thinkingConfig for Gemma models on the… — @chy1211

  • feat(xai): route xAI clients to Grok's native /v1/responses… — @ryanngit

  • feat(models): add a Settings → AI "Model Overrides" UI plus… — @xz-dev

  • feat(api): add Vary: Accept-Encoding to token-authenticated… — @chirag127

  • feat(sse): add Notion AI Web (Unofficial/Experimental)… (#6758)

  • feat(dashboard): add per-routing-combo compression-mode override to the… (#6760)

  • feat(resilience): operator-configurable account rotation policy — a new… — @artickc

  • feat(sse): preserve tools/tool_choice for tool-bearing… — @chirag127

  • chore(cursor): add Grok 4.5 effort/fast model IDs: (#6774 ). — @andrewmunsell

  • feat(codex): Codex provider model discovery now fetches the live… — @JxnLexn

  • feat(cursor): register the Opus 4.8, Fable 5, and Sonnet 5 model… — @andrewmunsell

  • Changelog fragments (changelog.d/): PRs now add their changelog entry as a new fragment…

  • feat(proxy): add a latency-optimized proxy rotation strategy that… — @iamraydoan

  • feat(db): include xp_audit_log in the automatic… (#6801)

  • feat(fusion): the fusion judge may now draw on its own knowledge and… — @chirag127

  • feat(dashboard): search box on the Playground's raw model <select> —… (#6811)

  • feat(usage): Antigravity/agy quota widget now surfaces the… (#4017 #6818)

  • feat(codex): Codex CLI compatibility shim — the Responses API… (#3697 #6820)

  • Z.ai Web (free web-session provider): new zai-web web-cookie provider drives the free… (#6823)

  • feat(dashboard): import multiple, possibly different… (#6836)

  • feat(compression): update the vendored GCF codec behind the… (#6837)

  • feat(sse): OpenRouter quota tracking — a dedicated fetcher polls… (#6842)

  • feat(sse): live gRPC-web quota fetcher for Grok Build (grok-cli)… (#6844)

  • feat(sse): add dual-window quota tracking for the… (#6845)

  • feat(sse): add a static local RPM budget (default… (#6846)

  • feat(sse): add quota tracking for the agentrouter… (#6850)

  • feat(providers): Add GPT-5.6 support across OpenAI API, Codex, and… (#6862) — @backryun

  • chore(providers): Align emitted Claude Code identity headers, bridge… (#6862) — @backryun

  • feat(api): add a structured X-Routing-Fallback-Reason… (#6872)

  • feat(api): new GET /api/usage/model-latency-stats management… (#6873)

  • feat(providers): add a vibeproxy-openai provider-node preset to `POST…

  • feat(usage): add… (#6875)

  • feat(sse): per-model default reasoning_effort… (#6879)

  • feat(providers): let a custom/openai-compatible connection opt into… — @andrea-kingautomation

  • feat(dashboard): add a Type filter (No Signup / OAuth Login / API Key)… (#6915)

  • feat(providers): expose an editable base-URL field on the ComfyUI… (#6928)

  • feat(providers): refresh the curated OpenRouter embeddings catalog… (#6976)

  • feat(quota): Add opt-in auto-ping to keep Codex quota windows warm —… (#6977)

  • feat(oauth): Add a one-click browser (PKCE) login for Grok Build… (#7013)

  • feat(sse): Add optional-enum null-omission idiom for… (#7023)

  • feat(auth): accept the x-goog-api-key header for client-facing… — @QRcode1337

  • feat(sse): add a local dual-window (5h + 7d,…

  • feat(api): opt-in scheduled auto-sync for free-proxy sources… — @chirag127

  • feat(kiro): register the GPT-5.6 Sol/Terra/Luna model family (272k… — @SemonCat

  • feat(dashboard): show the Codex subscription plan label in provider… — @CarmeloCampos

  • feat(dashboard): add a "Reorder" button to provider connections that… — @fzrilsh

  • feat(dashboard): add 180D and 365D periods to the Cost Explorer range… (#7213)

  • feat(sse): GitHub Copilot Claude models now route through… — @yidecode

  • feat(mitm): Antigravity MITM model mappings now support an optional… — @trfi

  • feat(sse): add native xAI Grok Imagine video generation provider —… — @anndev-69

  • feat(cli): add Grok Build CLI tool setup — writes a… — @rixzkiye

  • feat(provider): add Chenzk API OpenAI-compatible gateway. (thanks… — @CahyokPutraDev99

  • feat(dashboard): add a per-operator "hide this quota row" toggle to the… — @nguyenha935

  • feat(sse): session affinity (X-Session-Id / x-codex-session-id… — @tenshiak

  • feat(sse): gemini-web now honors the OpenAI tools array by… (#7286)

  • feat(eval): added a router-eval harness (npm run eval:router,… — @KooshaPari

  • feat(services): introduce a ServiceProviderPlugin contract… (#7333)

  • feat(routing): wire interceptFetch into the chat…

  • feat(dashboard): confirm before removing a single… (#7361)

  • feat(providers): Add a first-class xAI OAuth PKCE provider for… (#7399) — @fenix007

  • feat(dashboard): surface in-product guidance for the Settings → Prompt… (#7530)

  • feat(providers): Complete GHE Copilot OAuth provider implementation with… (#7546)

  • feat(providers): grok-web's add-connection dialog now… (#7567)

  • feat(providers): register lmstudio in the embedding provider registry… (#7601) — @ekinnee

  • feat(sse): honor a no-auth provider connection's… (#7622)

  • feat(dashboard): proxy-assignment UX quality-of-life fixes — the… — @tenshiak

  • feat(providers): zai-web (chat.z.ai) now wires live… — @andrea-kingautomation

  • feat(perplexity): Refresh Perplexity Web model mappings and Search API… (#7687) — @backryun

  • Generic OpenAI-compatible model sync now captures a… (#6879)

  • feat(dashboard): Kimi (Moonshot AI) official-partnership highlight on… (#7775)

  • feat(providers): let any provider connection — built-in… (#7818)

  • feat(routing): add a read-only `GET… (#7819)

  • feat(providers): copilot-m365-web tone-selected model… (#7872)

  • feat(shared): Add canonical numeric coercion helpers (toNumber,… (#7879)

  • feat(mcp): add a narrow mcp:connect API-key scope for… (#7895)

  • feat(dashboard): provider tab account search… (#7937)

  • feat(sse): classify grok-web Cloudflare anti-bot blocks… (#8019)

  • Teach the model the CCR retrieve protocol (marker →… (#8033)

  • Add a per-model/endpoint compression exclusion filter:… (#8034)

  • feat(i18n): zh-CN terminology glossary + consistency gate + 提供商→提供者… (#8038)

  • feat(dashboard): The Codex quota card no longer shows latent, never-used… (#8051)

  • feat(providers): native Fish Audio TTS provider on… (#8099)

  • feat(codex): add bounded multi-reference PNG/JPEG/WebP image editing… (#8122) — @xiaoyaner0201

  • refactor(sse): classify SSE critical-path empty catches… (#8142)

  • feat(db): persist caller session tag into call_logs for… (#8249)

  • feat(api): quota-aware fallback routing for web-fetch… (#8297)

  • feat(providers): map upstream reasoning-level metadata… (#8347)

  • feat(providers): weekly quota for xAI OAuth (Grok) (xai-oauth / xao)… (#8471) — @allanvb

  • feat(jobs): backup auto enable schedule is now… (#8513)

  • feat(ci): automate ratchet shrink-banking so file-size/complexity… (#8612) — @MumuTW

  • feat(providers): live monthly credit quota for Firecrawl (`GET… (#8759) — @allanvb

  • feat(search): add Firecrawl to POST /v1/search supported providers… (#8814) — @allanvb

  • feat(adobe-firefly): reference-image attach for generate + OpenAI…

  • Homologation suite: new npm run homolog runs the full…

  • feat(mcp): add a read-only Local Corpus context source with…

  • feat(providers): notion-web live model discovery via…

  • Provider connections can now be shown or hidden…

  • docs(i18n): rewrite Russian README as a full native…

  • Providers: adds the Xiaomi MiMo Token Plan provider and a… (#8861)

  • feat(opencode-plugin): auto-discover models while running + force sync (#8101) — @RaviTharuma

  • feat(combo): enforce provider and model family invariants (#8304) — @RaviTharuma

  • feat(cli): replace ANTHROPIC_SMALL_FAST_MODEL with Fable default (#8343) — @leszek3737

  • feat: read INITIAL_PASSWORD env var during setup (#8439) — @linhdmn

  • feat(providers): add missing opencode-go reasoning effort variants (#8441) — @Prudhvivuda

  • feat: add Claude Opus 5 support (#8464) — @backryun

  • feat(combos): add select all / unselect all in Browse Catalog (#8526) — @JoshimOfficial

  • feat(cli-tools): add all Hermes Agent auxiliary model roles (#8543) — @leszek3737

  • feat(ci): block stale UI translations when an English value is… (#8574)

  • feat(providers): flatten multi-turn history for gemini-web (#8607) — @Prudhvivuda

  • feat(sse): relay upstream 4xx error bodies verbatim on the… (#8622)

  • feat: Claude Code discovery aliases (surface non-Claude… (#8666)

  • feat(dashboard): copy-paste settings.json block for Claude Code discovery (#8722)

  • feat(quality): temporary relax of complexity/file-size ratchets for… (#8767)

  • feat(db): let the migration runner scan extra namespaced… (#8770)

  • feat(api): prompt-cache health summary endpoint and analytics tab (#8827)

⚡ Performance
  • perf(db): project columns + composite index in… (#6918) — @oyi77

  • perf(db): add jitter to stagger due-on-restart connections (#6919) — @oyi77

  • perf(startup): warm model catalog cache at module init (#6920) — @oyi77

  • perf(db): add temp_store=MEMORY pragma to SQLite init (#6921) — @oyi77

  • perf(db): cap modelLockouts eviction at 1000 entries (#6923) — @oyi77

  • perf: wrap ComboCard, HeroSection in React.memo (#7070) — @oyi77

  • perf: Date.now hoist, hasActiveDeltaValue hoist, buffer.split… (#7066) — @oyi77

  • perf(memory): mitigate event-loop starvation under 3000+ provider… (#7719) — @oyi77

  • perf: reduce long-context request copies (#7862) — @RaviTharuma

  • perf: lazy provider init, P2C quota cache, structuredClone… (#7893) — @oyi77

  • perf(api): singleflight version lookups (#8301) — @RaviTharuma

  • perf(api): skip the full catalog build for quota-exclusive keys (#8771)

🐛 Bug Fixes
  • fix: add re-entrancy guard to token health check sweep (#6917) — @oyi77

  • fix(grok): strip reasoningEffort for grok cli models (#6938) — @CitrusIce

  • fix(6954,6953): preserve system role + strip empty-signature thinking… (#6982) — @rafaumeu

  • fix(6980): classify Cloudflare AI neuron exhaustion as… (#6983) — @rafaumeu

  • fix(dashboard): hide disabled provider connections from combo builder (#6984)

  • fix(providers): cap grok-cli tools at 200 for cli-chat-proxy (#6986)

  • fix(6848): auto-cleanup for telemetry tables causing OOM (#6988) — @rafaumeu

  • fix(models): preserve direct-model combo metadata (#6993) — @JxnLexn

  • fix: DDG circuit breaker + null content validation (#7001) — @rafaumeu

  • fix(models): preserve chat-capable image model rows (#7004) — @xz-dev

  • fix(codex): preserve GPT-5.6 reasoning contract (#7012) — @xz-dev

  • fix(base-red): align least-used combo tests with executionKey usage… (#7015)

  • fix: infer bare models from active synced catalogs (#7028) — @guanbear

  • fix(auggie): update model registry to match v0.32.0 CLI model IDs (#7032) — @oyi77

  • fix(sse): register ollama-cloud in USAGE_FETCHER_PROVIDERS (#7041) — @alltomatos

  • fix(quality): read cognitiveComplexity= machine line in… (#7042) — @alltomatos

  • fix(providers): sanitize Claude native output_config.effort (#7050) — @xier2012

  • fix(combo): treat maxInputTokens as an input-only cap in the… (#7052) — @xier2012

  • fix(antigravity): collect native part.functionCall into tool calls (#7053) — @xier2012

  • fix(responses): map mid-conversation system turns to developer role (#7056) — @xier2012

  • fix(combo): least-used sorts by per-account executionKey (#7059) — @xier2012

  • fix(providers): AgentRouter model import applies Claude Code wire image… (#7060) — @xier2012

  • fix(translator): preserve thinking.budget_tokens: 0 in Claude->Gemini (#7061) — @xier2012

  • fix(cloudflare-relay): avoid invalid regex syntax in generated worker (#7063) — @SeaXen

  • fix(dashboard): strip browser-extension attrs before hydration (#7073) — @MrFadiAi

  • fix(relay): bound Bifrost stream lifetime (#7093) — @KooshaPari

  • fix(sse): recognize xiaomi-tokenplan mimo as a thinking-mode model (#7098)

  • fix(codex): strip regex lookaround from tool schema patterns (#7100)

  • fix(openai): strip reasoning_effort when GPT-5.x models carry… (#7101)

  • fix(compression): Headroom SmartCrusher skips developer-role messages… (#7102)

  • fix(providers): surface a warning on 404 model_not_found in… (#7103)

  • fix(executors): forward X-Session-ID/X-Title agent metadata headers (#7104)

  • fix(cli): verify better-sqlite3 native binary is actually loadable (#7105)

  • fix(sse): sanitize non-ok Antigravity streaming error body (port… (#7106)

  • fix(providers): add MiniMax image-generation provider (#7108)

  • fix(sse): handle space-separated arg name/value in Composer tool… (#7116)

  • fix(cli): remove MITM DNS spoof entries before killing server… (#7117)

  • fix(dashboard): include never-tested connections in combo builder… (#7118)

  • fix(api): check Vercel SSO-protection PATCH response on relay… (#7119)

  • fix(combos): reject oversized fusion panels before fan-out (port… (#7120)

  • fix(combo): detect empty content_block in streaming SSE peek (#7121)

  • fix(oauth): resolve Kiro AWS SSO cache client credentials by… (#7122)

  • fix(tests): vitest UI suite back to green (69 fails triaged — WS6.1) (#7127)

  • fix(auto): use p95 fallback in speed factors (#7128) — @KooshaPari

  • fix(models): update Anthropic model contextLength to 1M (#7129) — @HouMinXi

  • fix(ci): raise dast-smoke timeout 12->25min (build alone eats up… (#7139)

  • fix(compression): lazy-load typescript in RTK codeStripper so prod-lean… (#7164) — @alltomatos

  • fix(providers): accept m365.cloud.microsoft for copilot-m365-web token (#7166) — @xier2012

  • fix(executors): disable parallel tools for Codex Responses Lite (#7171) — @fenix007

  • fix(tests+providers): env-dependent tests exposed by GH-hosted runners (#6634… (#7174)

  • fix(combo): reject known context overflow without exhausting… (#7177) — @JxnLexn

  • fix: add static.cloudflareinsights.com to CSP script-src (#7178) — @oyi77

  • fix: extend turbopack ignoreIssue suppression to compression… (#7180)

  • fix: recognize Ollama Cloud session usage-limit 429 as… (#7181)

  • fix: preserve relayAuth for pool-referenced relay proxies (#7182)

  • fix: wire adaptive context-budget dial into settings schema… (#7183)

  • fix(providers): DuckDuckGo VQD 429 misclassified as 503 (#7185)

  • fix(db): cap OOM probe-failure cycle in getDbInstance() (#7186)

  • fix: stop opencode-go quota lookup defaulting to Z.AI… (#7187)

  • fix(providers): refresh OpenCode (oc) free-tier model catalog (#7188)

  • fix: include proxyId when testing a saved registry proxy (#7189)

  • fix: sanitize non-Latin1 chars in combo diagnostic headers (#7190)

  • fix: raise main server keepAliveTimeout/headersTimeout above… (#7191)

  • fix: route zai-web (and other registry-entry web-cookie… (#7192)

  • fix(providers): reject chat requests for cloud-agent-only jules provider (#7193)

  • fix: restore mobile grid-cols-1 fallback on quota page card… (#7194)

  • fix: wire modelAliases fetch into HermesAgentToolCard (#7195)

  • fix: surface real claude-web error body for non-SSE 400s (#7196)

  • fix(dashboard): agent bridge dns toggle uses POST, not PUT (#7197)

  • fix: stop duplicating text in Gemini Web streamed responses (#7198)

  • fix: filter hidden custom models out of legacy combo model… (#7199)

  • fix(dashboard): implement missing handleToggleSource on Free Pool tab (#7200)

  • fix: honor combo-level proxy assignments from the registry (#7201)

  • fix(ci): run quality gates on Mergify merge-queue draft PRs… (#7202)

  • fix: add dashboard-scoped typecheck gate covering… (#7203)

  • fix(guardrails/chat): stop Vision Bridge hijacking credentialed models to… (#7204) — @artickc

  • fix(translator): preserve Gemini thought parts as reasoning_content on… (#7206)

  • fix(translator): register openai response projection for gemini clients (#7207)

  • fix(cli): fast-path --version to skip full CLI bootstrap (#7208)

  • fix: honor PROVIDER_LIMITS_SYNC_SPACING_MS for local/API-key… (#7214)

  • fix(api): bulk-add API keys no longer overwrite existing… (#7234)

  • fix(sse): route the public OpenAI GPT-5.6 family through the… (#7242)

  • fix(providers): honor configured proxy on Grok Build egress (#7244)

  • fix(nvidia): expand NIM chat model catalog (#7247)

  • fix(sse): reconstruct Claude-format content in synthetic bypass… (#7248)

  • fix(build): isolate Windows HOME/AppData during next build (#7249)

  • fix(cli): omniroute dashboard respects PORT env when --port is… (#7252)

  • fix(sse): project non-streaming JSON back to the… (#7255)

  • fix(combo): fall back on Responses SSE failures (#7256) — @rushsinging

  • fix(routing): resolve nested combo-ref panel members in fusion… (#7259)

  • fix(usage): reset logs and show provider names in analytics (#7264) — @SeaXen

  • fix(sse): silence noisy proxy-failure log on caller-initiated… (#7266)

  • fix(codex): normalize nested Responses output content (#7269) — @JxnLexn

  • fix(combo): derive session stickiness key from Responses API… (#7277) — @alltomatos

  • fix(antigravity): wrap Pro fallback chain in try/catch for timeout… (#7290) — @HouMinXi

  • fix(logs): show saved provider names in request/provider log views (#7294) — @SeaXen

  • fix(stream): reconcile encrypted Codex reasoning visibility without… (#7304)

  • fix(build): packed tarball boot crash — server-ws timeout import… (#7308)

  • fix(skills): register cli-skill-collector in the agent-skills… (#7310)

  • fix(db): tolerate unavailable virtual table modules in stats (#7313) — @megamen32

  • fix(router-eval): retained-optimization gate cleanup (#7318) — @KooshaPari

  • fix(ci): Coverage job timeout 10->20min (lcov reporter pushed it… (#7342)

  • fix(electron): normalize hashed standalone externals (#7353) — @tianrking

  • fix(db): stop a 'latest' path segment from disabling backups and… (#7359) — @danscMax

  • fix(branding): regenerate raster favicons — white mark was shipped… (#7390) — @vzts

  • fix(test): skip real DNS writes in MITM dynamic-import test (#7398) — @HouMinXi

  • fix(antigravity): streaming passthrough for non-streaming clients (#7408) — @HouMinXi

  • fix(build): align engines.node with SUPPORTED_NODE_RANGE (#7490) — @alltomatos

  • fix(api): await params in Agent Bridge DNS route (Next.js 16) (#7492) — @alltomatos

  • fix(dashboard): show Obsidian context source card (#7500) — @DKotsyuba

  • fix(ci): fetch full base history in pr-test-policy (shallow… (#7501)

  • fix(sse): preserve chat quota across mixed windows (#7504) — @webmasterarbez

  • fix(oauth): surface sanitized device-code error instead of a… (#7511) — @danscMax

  • fix(oauth): repair qwen + codebuddy-cn device-code endpoints (#7517) — @danscMax

  • fix(mitm): strip trailing assistant prefill to prevent upstream… (#7520) — @chirag127

  • fix(codex): Test probe uses a ChatGPT-account-supported model (#7524)

  • fix(codex): validate refresh_token on import before persisting (#7525)

  • fix(codex): non-stream chat 502 'Response body is already used'… (#7526)

  • fix(oauth): surface tunnel hint when Codex OAuth runs on a remote… (#7527)

  • fix(sse): preserve custom tool output images (#7540) — @loulanyue

  • fix(combo): failover when upstream SSE is truncated mid-lifecycle (#7545) — @Chewji9875

  • fix(dashboard): prefer public endpoint URLs (#7547) — @nguyenha935

  • fix(cli): refresh runtime detection accurately (#7552) — @nguyenha935

  • fix(ui): improve React Flow dark theme (#7553) — @nguyenha935

  • fix(i18n): treat MISSING sync placeholders as absent in EN… (#7556)

  • fix(cli): Windows cert check/uninstall key off the real CA… (#7557)

  • fix(sse): feed compression pipeline the authoritative vision… (#7560)

  • fix(dashboard): providers model-name filter matches live/synced catalog (#7561)

  • fix(db): pre-init sql.js WASM ahead of any getDbInstance()… (#7562)

  • fix(dashboard): resolve costs page 500 from out-of-scope t() in… (#7564)

  • fix(dashboard): surface rate-limit warning on 429 chat-probe (#7565)

  • fix(sse): lazy-load playwright in claudeTurnstileSolver (#7566)

  • fix(sse): combo failover for OpenAI streams truncated without… (#7568)

  • fix(cli): reuse win32-aware locateCommand in tool-detector (#7569)

  • fix(codex): #7536 check content-type before touching response.body… (#7570)

  • fix(sse): stop dropping tool_search and leaking OpenAI-only… (#7571)

  • fix(api): resolve provider display name and dedup byModel on… (#7573)

  • fix(mitm): route Claude Code standalone MITM traffic (#7574) — @dongwook-chan

  • fix(sse): stop per-byte enumeration of binary image bytes in log… (#7576)

  • fix(sse): split effort/reasoning suffix off pinned cursor model… (#7577)

  • fix(chatgpt-web): recognize update_content.messages[] celsius WS frames (#7578)

  • fix(sse): 401 model-not-supported lockout + sticky… (#7580)

  • fix(antigravity): allow cloudcode envelope through messages guard (#7582) — @dongwook-chan

  • fix(sse): sanitize empty-signature thinking blocks + hoist… (#7583)

  • fix(sse): honor per-model targetFormat override for… (#7584)

  • fix(sse): clamp glm-4.6v max_tokens to the 32768 ceiling (#7585)

  • fix(cli): log Codex Responses WebSocket history/usage per logical… (#7588)

  • fix(providers): derive static model catalogs for search providers from… (#7589)

  • fix(stream-readiness): bump timeout for heavy Claude-format reasoning replicas (#7612) — @herjarsa

  • fix(translator): synthesize tool call chunks from response.completed… (#7613) — @ekinnee

  • fix(embeddings): add lmstudio to embedding provider registry (#7614) — @ekinnee

  • fix(combo): auto-clear stale session pins and emit recovery hints… (#7625) — @herjarsa

  • fix(providers): unify connection and routing flows (#7629) — @nguyenha935

  • fix(db): dedupe bulk-imported proxies by full credential tuple (#7644) — @alltomatos

  • fix(api): allow text-to-image on dual-modality models + revive… (#7648) — @danscMax

  • fix(stryker): add Microsoft Designer test to tap.testFiles (#7659)

  • fix(dashboard): cut UI import chain from connection persist module (CI… (#7677)

  • fix(perplexity-web): stop empty-content responses from live schematized SSE (#6955) — @artickc

  • fix(dashboard): make quota cards container responsive (#7027) — @xz-dev

  • fix(nvidia): restore GLM-5.2 reasoning on NIM (#7296) — @backryun

  • fix(i18n): complete Vietnamese dashboard localization and runtime… (#7493) — @nguyenha935

  • fix(providers): migrate muse-spark-web from GraphQL to WebSocket… (#7528) — @Ajeesh25353646

  • fix(combo): expose computed context_length via /api/combos for… (#7633) — @herjarsa

  • fix(api): enumerate tiered auto combo endpoints in… (#7662) — @ekinnee

  • fix(dashboard): topology reflects connection health + clears finished… (#7672) — @danscMax

  • fix(kimi-coding): capture and replay reasoning for thinking-mode turns (#7673) — @xz-dev

  • fix(ci): merge-train --fast mirrors test:unit subdir allowlist (#7688)

  • fix(sse): start credential-health sweep at boot so stale web… (#7689) — @danscMax

  • fix(cursor): discover models via official CLI command (#7692) — @makcimbx

  • fix(quota): fix antigravity/agy multi-model quota skipping in combos (#7695) — @irvandikky

  • fix(usage): preserve account identity history (#7700) — @xz-dev

  • fix(db): update proxies on password rotation (#7707) — @floze-the-genius

  • fix(providers): correct Chutes registry baseUrl (#7708)

  • fix(routing): strip prompt_cache_key for NVIDIA NIM (#7709)

  • fix(providers): degrade Arena (lmarena) cookie validation redirect to… (#7710)

  • fix(claude-web): unify Turnstile/executor/fast-path User-Agents behind… (#7711)

  • fix(sse): authenticate CLIProxyAPI fallback/passthrough legs with… (#7712)

  • fix(sse): proactively refresh Grok Build OAuth token before… (#7715)

  • fix(providers): classify ambiguous Mistral 401 instead of hard auth… (#7718)

  • fix(security): bump adm-zip >=0.6.0 + exact host matching in mitm DNS… (#7732)

  • fix(icons): fall back to Stepfun Mono when Color component is… (#7743) — @Dan-ex-hub

  • fix(stream): suppress </think> close marker for Responses API… (#7747) — @xz-dev

  • fix(sse): wire settings.wildcardAliases into model resolution (#7748)

  • fix(authz): classify forge/jcode CLI settings routes as LOCAL_ONLY (#7749)

  • fix(routing): honor eye-icon hidden models for no-auth providers in… (#7750)

  • fix(sse): persist rotated Gemini web-session cookies via… (#7751)

  • fix(docs): heal release-green docs drift + eslint any-suppression… (#7755)

  • fix(mcp): copy undici into dist/node_modules to prevent… (#7756)

  • fix(packaging): move fumadocs-mdx to devDependencies (#7757)

  • fix(ci): build API-only smoke workflows backend-only to fix… (#7758)

  • fix(cli): load DATA_DIR/server.env as fallback for .env on… (#7759)

  • fix(cli): split outboundUrlGuard's DB helpers so setup-opencode… (#7760)

  • fix(notion-web): production-ready labels, multi-workspace, inference,… (#7768) — @artickc

  • fix(kimi): expose K3 reasoning effort levels (#7776) — @xz-dev

  • fix(compression): apply compression combo assignments to routing combos (#7779) — @ekinnee

  • fix(i18n): regenerate Polish UI locale from English (#7782) — @leszek3737

  • fix(combo): retry transient errors in pipeline strategy (#7794) — @AndrianBalanescu

  • fix(quality): register nvidia-quota-phase1 and… (#7796)

  • fix(stream): synthesize terminal finish_reason chunk when upstream… (#7804) — @AndrianBalanescu

  • fix(plugins): 5 bugs on the plugin path (3 Windows-only, 2… (#7806) — @tmone

  • fix(cli): register ESM alias resolver for @/ paths under global… (#7808) — @rafaumeu

  • fix(auth): gate invalid-key check on isRequireApiKeyEnabled for… (#7810) — @AndrianBalanescu

  • fix(ci): repair release regressions exposed by clean runs (#7812) — @backryun

  • fix(rerank): add voyage format adapter for request/response… (#7813) — @AndrianBalanescu

  • fix(antigravity): attempt onboarding when projectId is empty (#5193… (#7815) — @rafaumeu

  • fix(stream): emit terminal SSE frames on mid-stream upstream failure (#7816) — @AndrianBalanescu

  • fix(sse): strip orphaned tool_use before antigravity/Vertex… (#7822)

  • fix(translator): sanitize tool_result.tool_use_id symmetrically with… (#7823)

  • fix(oauth): require chatgptUserId agreement for Codex account dedup (#7825)

  • fix(db): purge in-memory key-health state when a provider… (#7826)

  • fix(compression): keep a retrievable preamble instead of a bare CCR marker (#7827)

  • fix(db): log fatal boot-time SQLite driver-cascade failure… (#7828)

  • fix(docker): repair tls-client-node native binary after… (#7829)

  • fix(auth): restore TICK_MS in tokenHealthCheck (ReferenceError on… (#7830)

  • fix(cli): fix Windows CLI detection false negatives (#7753, #7774) (#7831)

  • fix(docs): document CREDENTIAL_REDACTION_ENABLED and… (#7833)

  • fix(dashboard): fix collapsed quota card session/weekly order (#7834)

  • fix(dashboard): mirror connection-row action-icon spacing under RTL (#7835)

  • fix: avoid cmd.exe spawn on Windows by using os.hostname()… (#7841) — @tientien17

  • fix(usage): harden account identity reconciliation (#7843) — @xz-dev

  • fix(cli): use rundll32 instead of cmd.exe for Windows browser… (#7844) — @tientien17

  • fix: add native lifecycle-aware health endpoint (#7852) — @RaviTharuma

  • fix: reserve chat admission before body parsing (#7853) — @RaviTharuma

  • fix: bound quadratic session-dedup memory growth (#7855) — @RaviTharuma

  • fix(compression): enable OmniGlyph for Claude Fable 5 (#7863) — @enjoyer-hub

  • fix(notion-web): add browser fingerprint headers to reduce Cloudflare… (#7864) — @HassiyYT

  • fix(mitm): gate Agent Bridge Repair on sudo password (#7865) — @skutanjir

  • fix(rerank): honor the connection's pinned proxy on rerank calls (#7867)

  • fix(compression): skip CCR on tool outputs to preserve agent loop (#7869) — @herjarsa

  • fix(vision-bridge): reroute auto/ prefix to vision model when images present (#7871) — @herjarsa

  • fix(opencode-plugin): support separate management read token (#7885) — @RaviTharuma

  • fix(sse): CC bridge loses OpenAI-format image input… (#7888)

  • fix(combo): strip boolean reasoning field for opencode-go providers (#7891) — @AndrianBalanescu

  • fix(notion-web): accept OpenAI content-parts arrays in transcript (#7896) — @artickc

  • fix(notion-web): reuse threadId across OpenAI multi-turn (no new chat… (#7900) — @artickc

  • fix(gemini): strip OpenAI "strict" tool-schema keyword for… (#7901) — @Witroch4

  • fix(antigravity): collect native functionCall parts in SSE collector (#7902) — @Witroch4

  • fix(sse): recover invalid Anthropic thinking signatures once (#7906) — @insoln

  • fix(resilience): don't cool down accounts or trip the breaker on client… (#7908) — @insoln

  • fix(dashboard): preserve quota cutoff drafts (#7909) — @hydraxman

  • fix(providers): treat public-host 302 as valid in Gemini Web connection… (#7917)

  • fix(providers): read reasoning_text in Claude-format response translator (#7919)

  • fix(dashboard): safely render structured error objects in Request Logs… (#7920)

  • fix(dashboard): repair monaco deep import broken by 0.56 exports map (#7922)

  • fix(autostart): adopt 9Router VBS startup to suppress console flash on… (#7925) — @tientien17

  • fix(translators): normalize TitleCase tool names for non-Anthropic models (#7926) — @nramabad

  • fix(api): resolve local provider models via dashboard catalog… (#7927) — @ekinnee

  • fix(auto): pool accounts by provider model (#7928) — @adrianaryaputra

  • fix(perplexity-web): multi-step empty content + advanced-quota cooldown (#7930) — @artickc

  • fix(ccr): resolve principal via OMNIROUTE_API_KEY env var on… (#7932) — @ekinnee

  • fix(combo): context-aware fallback ignores model_context_override (#7933) — @tmone

  • fix(i18n): preserve remaining Vietnamese localization (#7935) — @nguyenha935

  • fix(mitm): gate Agent Bridge DNS and Trust Cert on sudo password (#7939) — @skutanjir

  • fix(providers): route iflytek/sparkdesk to Spark's OpenAI-compatible… (#7942) — @FenjuFu

  • fix(sse): preserve parallel_tool_calls for GPT-5.6 delegation… (#7957)

  • fix(providers): copilot-m365-web fails loudly on empty turns +… (#7958)

  • fix(providers): treat unreliable web-cookie /models probe status as… (#7959)

  • fix(api): add amazon-q to the static model catalog (#7960)

  • fix: parse Gemini 429 RetryInfo.retryDelay for model lockout (#7961)

  • fix(quality): tolerate ESLint's trailing unpruned-suppressions text… (#7962)

  • fix(cli): translate missing sqlite bindings error into actionable… (#7963)

  • fix(cli): spawn opencode.cmd shim with shell:true on win32 (#7964)

  • fix(dashboard): correct block-extra-Claude-usage toggle copy to match… (#7965)

  • fix(api): classify /api/acp/agents as loopback-only (#7966)

  • fix(auth): restore configurable HEALTHCHECK_BATCH_SIZE dropped by… (#7970)

  • fix(test): widen ratelimit-admission pollUntil deadline to 10s (#7971)

  • fix(pricing): clarify disabled automatic sync status (#7972) — @RaviTharuma

  • fix(combo): exempt content_filter from empty-content detection (#7973) — @HouMinXi

  • fix(embeddings): support secure multimodal inputs (#7978) — @RaviTharuma

  • fix(resilience): cap exactCooldownMs against maxCooldownMs (#7980) — @ekinnee

  • fix(electron): derive macOS Helper name from execPath to remove 2nd… (#8002)

  • fix(chatcore): report string-reason client aborts as 499, not 502 (#8011) — @Long-Feeds

  • fix(models): drop generic catalog siblings of specialty surfaces (#8021) — @RaviTharuma

  • fix(models): stop inventing chat capabilities for specialty surfaces (#8022) — @RaviTharuma

  • fix(capabilities): resolve models.dev specialty rows across provider keys (#8023) — @RaviTharuma

  • fix(models): attach models.dev pricing to GET /v1/models entries (#8025) — @RaviTharuma

  • fix(grok-cli): require full auth.json on OAuth paste import (#8027) — @RaviTharuma

  • fix(grok-cli): sanitize function_call_output before Grok Build dispatch (#8030) — @RaviTharuma

  • fix(sse): bound forwarded response headers (#8041) — @insoln

  • fix(sse): replace spoofable .includes() PromptQL issuer check… (#8042)

  • fix(sse): bound Codex SSE peek read with per-read timeout (#8043)

  • fix(cli): stop double-prefixing combo model ids in opencode… (#8047)

  • fix(antigravity): scope 404 model-not-found lockout to exact model +… (#8050) — @AndrianBalanescu

  • fix: repair pre-existing red gates on the release/v3.8.49 tip (#8055)

  • fix(oauth): honor connectionId on token refresh so email-less… (#8062) — @insoln

  • fix: classify Google quota exhaustion responses (#8071) — @rafaumeu

  • fix(providers): refresh duckduckgo-web catalog to current Duck.ai wire… (#8079)

  • fix(security): decouple request PII redaction from injection mode (#8102) — @RaviTharuma

  • fix(providers): discover live AGY models (#8123) — @adevwithpurpose

  • fix(guardrails): align INPUT_SANITIZER request masking gate (#8124) — @RaviTharuma

  • fix(providers): refresh Baidu ERNIE and Qianfan website URLs (#8128) — @TrackCrewGalore

  • fix(stream): add logging to empty catch blocks in stream error… (#8143) — @chirag127

  • fix(sse): stop Codex/Responses sanitizer turning system image_url… (#8147)

  • fix(db): register SIGHUP handler and stop force-killing server… (#8148)

  • fix(providers): add missing poe registry baseUrl entry (#8149)

  • fix(routing): anchor quota cache on globalThis for cross-chunk… (#8150)

  • fix(responses): close namespace round-trip for Responses-Chat… (#8151) — @RCrushMe

  • fix(oauth): warn instead of silently opening unreachable localhost… (#8152)

  • fix(db): stop closing the sql.js singleton in getDbInstance()… (#8153)

  • fix(sse): run compression pipeline per turn in Codex Responses WS… (#8154)

  • fix(cli): merge node bin dir into CLI healthcheck PATH for codex… (#8156)

  • fix(sse): anonymous fingerprint fallback for keyless Pollinations… (#8157)

  • fix(cli): surface the real spawn error in process supervisor (#8158)

  • fix: strip internal reasoning placeholder from user-visible… (#8162) — @Dingding-leo

  • fix(combos): expose synced reasoning-effort variants in Combo… (#8165) — @Dingding-leo

  • fix(windows): add windowsHide to all child process spawns (#8167) — @Dingding-leo

  • fix(cursor): bridge native tools to client calls (#8171) — @makcimbx

  • fix(security): bound JWT-extraction regexes to prevent polynomial… (#8173)

  • fix: log pending request counter decrement failures (#8179) — @rafaumeu

  • fix: suppress sql.js build warning via non-analyzable… (#8184) — @rafaumeu

  • fix: align INPUT_SANITIZER_ENABLED default to true across… (#8185) — @rafaumeu

  • fix(combo): skip remaining same-provider targets on 401/403 auth… (#8195) — @rafaumeu

  • fix(compression): make memo key model-independent for non-vision engines (#8196) — @rafaumeu

  • fix(resilience): add max/step to NumberField for provider cooldown inputs (#8203) — @rafaumeu

  • fix(providers): fix Azure AI Foundry multi-model discovery and… (#8206) — @not-knope

  • fix(logs): stop the async-EPIPE log-flood loop at its ignition… (#8207) — @Tasogarre

  • fix(sse): surface OpenRouter mid-stream error chunks instead of a… (#8210) — @hartmark

  • fix(sse): Gemini malformed function-call handling + tool_choice… (#8211) — @hartmark

  • fix(sse): tool-incapable provider handling (AI Horde + Responses… (#8212) — @hartmark

  • fix(sse): Gemini TPM/RPD quota classification + combo… (#8213) — @hartmark

  • fix(services): resolve and record a real pid when adopting a service (#8218) — @seanford

  • fix(memory): resolve remote embedding dimensions for reindex (#8220) — @Prudhvivuda

  • fix(dashboard): correct machine-translated Korean UI strings in ko.json (#8224) — @MichaelYcJo

  • fix(devin-cli): refresh shared model catalog (#8227) — @backryun

  • fix(claude-web): align session transport and fallback (#8230) — @backryun

  • fix: restore OAuth auto-refresh for gemini-cli connections (#8232) — @seanford

  • fix: normalize Codex URLs and dashboard regressions (#8233) — @nguyenha935

  • fix(api): narrow claudeClassifierCompat auto trigger so… (#8236)

  • fix(gemini): drop HARM_CATEGORY_CIVIC_INTEGRITY from the default… (#8238)

  • fix(backend): word-boundary-safe tool-result truncation in lite… (#8239)

  • fix(providers): filter unsupported family-fallback candidates against… (#8240)

  • fix(translator): synthesize tool call chunks from response.completed… (#180 #3980)

  • fix(api): Vercel Relay deploy now checks the Deployment… (#1037) — @ricatix

  • fix(oauth): resolve Kiro AWS SSO cache client credentials by… — @XCrag

  • fix(combo): streaming Claude responses whose content block opens… (#1382) — @heishen6

  • fix(codex): strip regex pattern lookaround (lookahead/lookbehind)… (#7100) — @evinjohnn

  • fix(cli): stopMitm() now removes /etc/hosts DNS-spoof entries… — @dionisius95

  • fix(sse): Cursor Composer/Auto tool calls that separate the arg… — @way-art

  • fix(dashboard): the "Custom Models" add/edit form now has a "Vision… — @nguyenphi37

  • fix(combos): fusion combos now reject an oversized panel (>40 models… — @fontvu

  • fix(providers): the OpenAI-compatible "Check" validation flow now… — @advane204f

  • fix(dashboard): include never-tested custom provider connections in the… — @fajarbossit

  • fix(executors): forward agent-supplied X-Session-ID/X-Title… (#7104) — @chitholian

  • fix(sse): Antigravity streaming requests that hit a non-ok… — @Duongkhanhtool

  • fix(providers): MiniMax Text-to-Image now works — a minimax… — @felipeleite

  • fix(cli): the runtime self-heal now verifies a cached… (#2493) — @mrprohack

  • fix(openai): strip reasoning_effort/reasoning for GPT-5.x models… — @techsolutionmta

  • fix(providers): preserve relayAuth for… (#5716)

  • fix(providers): wire the Devin cloud-agent provider… (#6142)

  • fix(providers): refresh Baidu ERNIE + Qianfan provider website URLs to… (#6271) — @TrackCrewGalore

  • fix(providers): honor a provider-level proxy assigned… (#6272)

  • fix(api): merge tool_call continuation deltas that… (#6276)

  • fix(providers): modernize the lmarena provider for the Arena.ai… (#6280) — @backryun

  • fix(providers): web-provider model discovery updated — qwen-web uses… — @janeza2

  • fix(logs): the request-log detail modal no longer reopens by… — @xz-dev

  • fix(providers): update SenseNova Token Plan support — register the… — @xz-dev

  • fix(providers): give v0-vercel-web its own alias so its… (#6343)

  • fix(providers): route AgentRouter key validation… (#6377)

  • fix(db): stop legacy log-archive migration from…

  • fix(docs): document Turbopack build memory tradeoff and… (#6409)

  • fix(compression): surface silently-dropped…

  • fix(providers): honor the max_token capability… (#6524)

  • fix(dashboard): the onboarding tier-flow diagram rendered broken — its… — @ianriizky

  • fix(routing): the auto combo's no-auth candidate pool now honors a… (#6557)

  • fix(sse): server-tool literal names (e.g. web_search) are… — @MikeTuev

  • fix(sse): sanitize non-Latin1 characters before… (#6612)

  • fix(api): recognize OpenRouter… (#6623)

  • fix(db): share one in-flight sql.js load across… (#6628)

  • fix(db): unwrap lone named-parameter objects before… (#6802)

  • fix(db): break probe-failed/restore loop on large storage.sqlite: (#6632 ). — @KooshaPari

  • fix(ci): exclude check-test-masking.test.ts's own… (#6634)

  • fix(routing): recognize Kimi-style "exceeded model… (#6637)

  • fix(cli): Claude Code installed via WinGet is now detected on… — @enjoyer-hub

  • fix(providers): removed obsolete/defunct providers from the catalog… — @backryun

  • fix(sse): requests rejected before handleChatCore… (#6698)

  • fix(providers): reject chat-completions requests for… (#6699)

  • fix(sse): unwrap bare {function:{…}} tools so OpenAI-shape… — @samir-abis

  • fix(oauth): stop merging distinct Codex OAuth logins that share an… — @lucasjustinudin

  • fix(codex): detect "model at capacity"/overloaded errors embedded… — @ryanngit

  • fix(volcengine): clamp max_tokens to the VolcEngine Ark endpoint cap… — @whale9820

  • fix(antigravity): surface aborted/malformed Gemini tool calls (e.g.… — @anhdiepmmk

  • fix(routing): the reasoning-token headroom buffer clamps to the… (#6714) — @xz-dev

  • fix(api): omniroute health (and health components/`health… (#6677 #6717)

  • fix(startup): webpack build broke on case-insensitive filesystems… (#6718)

  • fix(build): Turbopack production build emitted an "Overly broad… (#6582 #6720)

  • fix(providers): Codex Desktop requests to gpt-5.3-codex-spark failed… (#6651 #6721) — @alltomatos

  • fix(providers): the provider quota card's weekly/session bars re-sorted… (#6687 #6722)

  • fix(i18n): pt-BR was missing 194 UI keys present in en.json — a… (#6695 #6723)

  • fix(startup): omniroute --mcp crashed at Node ESM link time with… (#6559 #6725)

  • fix(providers): Kiro sent the adaptive-thinking… (#6576 #6726)

  • fix(translator): Cursor's local Subagent tool call is no longer… — @like3213934360-lab

  • fix(translator): GLM 5.2 (and other OpenAI-compatible upstreams that… — @itiwant

  • fix(resilience): OmniRoute didn't respect an exhausted Ollama Cloud (or… (#6638 #6731)

  • fix(resilience): a combo step "pinned" to one fingerprint account… (#6696 #6732)

  • fix(api): Responses passthrough emitted event-only SSE frames (no… (#6561 #6735)

  • fix(compression): /api/compression/preview's top-level… (#6488 #6741)

  • fix(resilience): account selection could pick an account already out of… (#6686 #6742)

  • fix(api): reasoning_content (extended-thinking text) was… (#6662 #6743)

  • fix(api): the compression config PUT schema now accepts… — @alltomatos

  • fix(api): raised the provider apiKey length cap for… — @alltomatos

  • fix(routing): fusion combos no longer silently drop combo-ref panel… (#6764)

  • fix(ci): publish electron-updater latest*.yml… (#6766)

  • fix(i18n): translate hardcoded Portuguese dashboard strings to English (#6761, #6768): (#6769 ). — @chirag127

  • fix(providers): strip redundant node prefix when… (#6772)

  • fix(providers): scope nvidia NIM 404s to the single… (#6773)

  • fix(codex): bump the default Codex CLI client identity from… (#6780) — @quanturbo

  • fix(ci): the blocking "Impacted unit tests (TIA)" step… (#6788)

  • fix(translator): read PDF/video file_data attachments on the… — @Witroch4

  • fix(providers): ensure DeepSeek Web SSE emits [DONE] after FINISHED: (#6791 ). — @Pitchfork-and-Torch

  • fix(api): the compression config PUT schema… — @Pitchfork-and-Torch

  • fix(electron): materialize Turbopack hashed-module symlinks during packaging (#6724, #6594): (#6794 ). — @huohua-dev

  • fix(cursor): send the Agent CLI build id as… — @andrewmunsell

  • fix(cli): waitForServer() no longer reports ready from… (#6800)

  • fix(sse): de-flake timing-sensitive combo… (#6803)

  • fix(codex): strip include from compact responses requests: (#6805 ). — @yinaoxiong

  • fix(dashboard): surface Claude extraUsage credits in… (#6806)

  • Request count by provider & date: Dashboard → Analytics now shows a dedicated table of… (#6812) — @tjengbudi

  • fix(resilience): an Ollama Cloud (or any apikey-category provider)… (#3709 #6817)

  • fix(providers): an explicit thinking.budget_tokens: 0 is now honored… — @alltomatos

  • fix(bootstrap): filter empty process.env values before spawning… — @AndrianBalanescu

  • fix(providers): classify upstream 404 responses as MODEL_NOT_FOUND… — @AndrianBalanescu

  • fix(db): cap the sql.js OOM-during-probe path in… (#6632 #6835)

  • fix(mcp): de-duplicate TOTAL_MCP_TOOL_COUNT by tool… (#6854)

  • fix(usage): xAI's exact provider-reported cost_in_usd_ticks no… (#6856) — @KooshaPari

  • fix(plugin): the @omniroute/opencode-plugin dynamic provider hook… (#6859)

  • fix(sse): apply cliproxyapiModelMapping at CLIProxyAPI… (#6876)

  • fix(sse): defer response.completed until a trailing… (#6906)

  • fix(cli): ship head-response-guard.cjs into the standalone… (#6908)

  • fix(sse): wire the shared quota-fetch throttle into… (#6911)

  • fix(sse): rename client-sent max_completion_tokens to… (#6912)

  • fix(providers): PROVIDER_LIMITS_SYNC_SPACING_MS now… (#6916)

  • fix(sse): classify LAN embeddings providers (10/8,… (#6925)

  • fix(sse): Qwen Web executor no longer sends `[object… (#6927)

  • perf(api): relay chat-completions routes now thread the… (#6930)

  • fix(sse): normalizeCodexMessageContentPart now… (#6932)

  • fix(sse): omit removed attachments field from Muse… (#6935)

  • fix(dashboard): label audio/embeddings/image compatible… (#6936)

  • fix(api): model-list discovery for LAN-local… (#6939)

  • fix(providers): openai->gemini transform now maps… (#4170) — @rafaumeu

  • fix(sse): escape backslash before brackets in ChatGPT-web… (#6944) — @brick30llc-ctrl

  • fix(oauth): tokenHealthCheck now lowercase-normalizes… (#6947)

  • fix(sse): make Codex Responses tool-arg normalization… (#6951)

  • fix(sse): drop internal commentary-phase Responses… (#6952)

  • fix(sse): stop forwarding empty-signature thinking… (#6953)

  • fix(sse): perplexity-web's default pplx-auto/pplx-sonar… — @artickc

  • fix(combos): embeddings-only and rerank-only models (e.g. JinaAI,… (#6975)

  • fix(combos): when 2+ distinct model ids from the same provider would… (#6957)

  • fix(dashboard): the combos builder now hides provider connections the… — @attid

  • fix(providers): cap grok-cli tools at 200 per request, matching xAI's… — @gitcommit90

  • fix(providers): DuckDuckGo AI Chat executor propagates… (#6996)

  • fix(providers): refresh OpenCode (oc) free-tier model… (#6998)

  • fix(api): raise the main server's… (#7003)

  • fix(compression): wire the adaptive context-budget… (#7005)

  • fix(usage): stop opencode-go quota lookup from… (#7022)

  • fix(providers): Auggie (Augment CLI) model registry updated to the real… — @oyi77

  • fix(ci): add a dashboard-scoped typecheck gate covering… (#7033)

  • fix(cli): omniroute dashboard (no --port flag) now respects… — @kaon0388v1

  • fix(build): extend the Turbopack ignoreIssue… (#7051)

  • fix(providers): web-cookie… (#7058)

  • fix(resilience): recognize Ollama Cloud's 5-hour… (#7071)

  • fix(dashboard): restore mobile single-column fallback… (#7072)

  • fix(dashboard): include proxyId when testing a saved… (#7080)

  • fix(sse): xiaomi-tokenplan mimo models (e.g. mimo-v2.5-pro)… (#7098) — @xxue-z

  • fix(dashboard): align onboarding tier descriptions and localize the… (#7125) — @Wibias

  • fix(sse): claude-web now surfaces the real upstream error body… (#7134)

  • fix(db): honor combo-level proxy assignments from the… (#7149)

  • fix(dashboard): wire modelAliases fetch into… (#7151)

  • fix(dashboard): filter hidden custom models out of the… (#7156)

  • fix(dashboard): Agent Bridge DNS toggle now sends POST… (#7157)

  • fix(dashboard): implement missing handleToggleSource… (#7161)

  • fix(sse): stop duplicating text in Gemini Web streamed… (#7163)

  • fix(executors): Codex Responses Lite requests force serial tool calls… (#7171) — @fenix007

  • fix(translator): preserve Gemini thinking-mode thought:true parts as… — @warelik

  • fix(translator): register the missing OpenAI→Gemini response projection… — @warelik

  • fix(cli): omniroute --version now fast-paths before the tsx/esm… — @Jordannst

  • fix(ci): build dast-smoke and nightly API-only smoke… (#7226)

  • api: bulk-add API keys no longer overwrite existing provider… — @asynx6

  • fix(sse): feed the compression pipeline the… (#7237)

  • fix(sse): route the public OpenAI GPT-5.6 family (gpt-5.6,… — @Jordannst

  • fix(providers): honor a configured proxy on Grok Build egress — the… — @ryanngit

  • fix(nvidia): expand NIM chat model catalog with newly-observed… — @spacesky-cell

  • fix(sse): synthetic bypass responses for Claude-format clients no… — @KunN-21

  • fix(build): isolate Windows HOME/AppData during next build. (thanks… — @KunN-21

  • fix(dashboard): providers model-name filter now matches… (#7250)

  • fix(docs): correct stale /api/version and… (#7253)

  • fix(sse): project non-streaming JSON responses back to… (#7255) — @warelik

  • fix(i18n): treat __MISSING__: sync-script… (#7258)

  • fix(authz): classify /api/cli-tools/forge-settings and… (#7263)

  • fix(sse): lazy-load playwright in claudeTurnstileSolver… (#7265)

  • fix(sse): stop logging a caller-initiated request abort/timeout… — @TuyulSpam

  • fix(sse): classify 401 "model X is not supported" as… (#7268)

  • fix(dashboard): resolve `ReferenceError: t is not… (#7272)

  • fix(cli): Windows MITM root-CA check/uninstall keyed… (#7275)

  • fix(cli): reuse cliRuntime's win32-aware… (#7279)

  • fix(dashboard): connection Test surfaces a rate-limit… (#7284)

  • fix(sse): combo failover now detects OpenAI-shape… (#7285)

  • fix(db): getDbInstance() now guarantees sql.js WASM has…

  • fix(sse): split effort/reasoning suffix off pinned… (#7289)

  • fix(sse): hoist client-injected system messages to… (#7293)

  • fix(sse): treat Uint8Array/Buffer as opaque binary in… (#7297)

  • fix(cli): load DATA_DIR/server.env as a fallback for… (#7302)

  • fix(rerank): Honor the connection's pinned proxy on rerank calls, so… (#7350) — @kamenkadmitry

  • fix(electron): Normalize hashed Turbopack external imports in packaged… (#7353) — @tianrking

  • fix(chatgpt-web): recognize update_content.messages[]… (#7357)

  • fix(sse): clamp glm-4.6v max_tokens to the 32768… (#7364)

  • fix(sse): honor per-model targetFormat override for… (#7364)

  • fix(sse): combo session stickiness now releases a… (#7387)

  • fix(cli): log Codex Responses WebSocket history/usage… (#7388)

  • fix(embeddings): remove the non-existent… — @kamenkadmitry

  • fix(providers): expose a base-URL override for… (#7447)

  • fix(build): align engines.node supported range… (#7446)

  • fix(i18n): replace machine-bulk-filled Vietnamese UI… (#7493)

  • fix(db): stop closing the single sql.js singleton… (#7494)

  • fix(sse): stop stream readiness from treating a… (#7503)

  • Fixed the Codex connection Test button always… (#7521)

  • The Codex account import (`POST… (#7522)

  • The PKCE OAuth start… (#7523)

  • fix(providers): search providers now expose a static… (#7529)

  • fix(sse): map tool_search to a Chat function tool… (#7532)

  • fix(sse): gate verbosity/prompt_cache_key on OpenAI… (#7533)

  • fix(api): Usage page "by provider" table now shows the… (#7534)

  • fix(api): Usage page "model usage" table no longer… (#7535)

  • fix(codex): non-stream Codex (ChatGPT-account) chat no… (#7536)

  • fix(sse): preserve input_image parts in array-valued Codex… (#7540) — @loulanyue

  • fix(providers): stop reporting Arena (lmarena) cookie… (#7542)

  • fix(combo): streaming combo failover now fails over when an… (#7545) — @Chewji9875

  • fix(dashboard): Public and managed tunnel endpoints now take precedence… (#7547) — @nguyenha935

  • fix(claude-web): unify Turnstile solver, executor and… (#7548)

  • fix(cli): fall back to the node:sqlite driver cascade… (#7586)

  • fix(api): expose Responses-API-format (OpenAI/Codex)… (#7587)

  • grok-cli: OAuth paste-import now requires the full… (#7610)

  • fix(sse): proactively refresh Grok Build's expiring… (#7610)

  • grok-cli: sanitize Responses function_call_output.output values… (#7611)

  • fix(translator): guard against double-emission when response.completed

  • fix(routing): strip prompt_cache_key for NVIDIA NIM —… (#7617)

  • fix(routing): honor eye-icon hidden models for no-auth… (#7620)

  • fix(providers): correct Chutes registry baseUrl from… (#7621)

  • fix(providers): unify connection status across… (#7629)

  • fix(providers): classify Mistral ambiguous 401 (quota… (#7638)

  • fix(sse): route CLIProxyAPI fallback/passthrough legs… (#7645)

  • fix(packaging): move fumadocs-mdx from dependencies to… (#7661)

  • GET /api/combos/auto now enumerates template variants…

  • fix(sse): persist rotated Gemini web-session cookies… (#7676)

  • fix(dashboard): mirror connection-row action-icon… (#7680)

  • fix(cli): split outboundUrlGuard.ts's DB/feature-flag… (#7682)

  • fix(sse): wire settings.wildcardAliases into model… (#7693)

  • fix(mcp): copy undici into dist/node_modules to prevent… (#7701)

  • fix(translator): sanitize tool_result.tool_use_id symmetrically with… (#7705)

  • fix(db): bulk proxy imports now update an existing proxy when… (#7707) — @floze-the-genius

  • fix(oauth): stop Codex OAuth completion from collapsing… (#7737)

  • fix(db): purge in-memory API-key health/rotation state… (#7740)

  • perf(db): the 5 provider-connection/node call sites 's IC2… (#7787 #7744) — @oyi77

  • fix(compression): CCR no longer collapses a whole… (#7746)

  • fix(sse): strip orphaned tool_use blocks before… (#7752)

  • fix(cli): resolve Windows CLI detection false negatives…

  • fix(dashboard): keep session/weekly quota rows in a… (#7764)

  • fix(db): log the fatal boot-time SQLite driver-cascade… (#7773)

  • fix(sse): Claude-Code-compatible bridge (AgentRouter and any CC… (#7777) — @beingshafin

  • perf(db): getProviderConnections now reads from a shared… (#7787) — @oyi77

  • fix(docs): document CREDENTIAL_REDACTION_ENABLED and… (#7793)

  • fix(combo): retry intermediate pipeline-strategy steps… — @AndrianBalanescu

  • fix(docker): repair tls-client-node native binary after… (#7802)

  • fix(runtime): Apply refreshed request-queue settings, restore… (#7812)

  • fix(oauth): attempt Antigravity onboarding inline when… (#7814)

  • fix(api): add amazon-q to the static model catalog so… (#7820)

  • fix(sse): preserve parallel_tool_calls for GPT-5.6… (#7821)

  • fix(quality): tolerate ESLint's trailing… (#7837)

  • fix(test): widen ratelimit-admission pollUntil deadline… (#7842)

  • fix(cli): Windows omniroute dashboard browser fallback uses… (#7844)

  • fix(dashboard): Request Logs detail modal no longer… (#7845)

  • fix(backend): stop buildClientRawRequest deep-cloning… (#7847)

  • fix(sse): copy the combo attempt body shallowly instead… (#7847)

  • fix(sse): estimate the combo fallback-compression… (#7847)

  • fix(providers): Copilot reasoning_text now surfaces in… (#7856)

  • fix(providers): treat a non-401/403 response from the… (#7857)

  • fix(providers): copilot-m365-web now fails loudly on a…

  • fix(providers): Gemini Web connection test now accepts… (#7859)

  • fix(cli): translate missing better-sqlite3 native… (#7868)

  • fix(auth): restore HEALTHCHECK_BATCH_SIZE env-var… (#7875)

  • fix(opencode-plugin): the official @omniroute/opencode-plugin now accepts… (#7884) — @RaviTharuma

  • fix(dashboard): Provider Quota cutoff inputs preserve unsaved values… (#7889)

  • fix(sse): Anthropic requests that reject a completed historical… (#7906)

  • Chat Core: client aborts that reject the upstream fetch with a raw…

  • fix(resilience): Client-aborted requests are now treated as client… (#7907)

  • fix(cli): spawn the win32 opencode.cmd shim with… (#7913)

  • fix(dashboard): correct "Block extra Claude usage"… (#7918)

  • fix(autostart): Windows auto-start no longer flashes a console window —… (#7925)

  • fix(providers): perplexity-web now detects the… (#7930) — @artickc

  • fix(sse): parse Gemini 429 RetryInfo.retryDelay /… (#7940)

  • fix(electron): resolve macOS Helper by binary name to… (#7941)

  • fix(providers): iFlytek Spark (iflytek, sparkdesk) now target the… (#7942) — @FenjuFu

  • fix(api): classify /api/acp/agents as loopback-only… (#7948)

  • fix(cli): stop double-prefixing combo model ids in the… (#7976)

  • fix(providers): route noauth opencode-zen connections… (#7993)

  • fix(providers): refresh duckduckgo-web free model… (#8000)

  • fix(providers): repoint the zai-web executor at… (#8014)

  • fix(sse): bound the Codex SSE peek/passthrough body… (#8020)

  • fix(sse): replace spoofable .includes() PromptQL… (#8029)

  • fix(providers): path-shaped multimodal model ids (e.g.… (#8032) — @Prudhvivuda

  • fix(cli): merge Node's own bin dir into the CLI… (#8036)

  • fix(sse): Bound forwarded upstream response headers so large… (#8041) — @insoln

  • fix(db): register SIGHUP handler and stop force-killing… (#8045)

  • fix(oauth): warn instead of silently opening an… (#8046)

  • fix(sse): run the prompt-compression pipeline per turn in the… (#8052)

  • fix(oauth): Refreshing the token on an email-less OAuth connection… (#8059)

  • fix(routing): anchor quota routing cache on globalThis… (#8065)

  • fix(providers): add the missing poe REGISTRY entry so built-in Poe… (#8082)

  • fix(translator): set status: "completed" on… (#8083)

  • fix(sse): give keyless Pollinations image requests the… (#8085)

  • fix(sse): stop Codex/Responses sanitizer from turning… (#8089)

  • fix(cli): surface the real spawn error… (#8091)

  • fix(antigravity): preserve request/tool fidelity, make credits modes… (#8098) — @nguyenha935

  • fix(dashboard): clamp provider-cooldown min/max ms to… (#8107)

  • fix(providers): filter unsupported family-fallback… (#8134)

  • fix(providers): classify HTTP 400 model-unavailable as… (#8136)

  • fix(backend): word-boundary-safe tool-result truncation… (#8169)

  • fix(api): narrow claudeClassifierCompat auto trigger so… (#8189)

  • fix(providers): carve cookie-auth providers out of… (#8200)

  • fix(gemini): drop HARM_CATEGORY_CIVIC_INTEGRITY from… (#8231)

  • fix(dashboard): Normalize Codex client URLs, restore the analytics… (#8233) — @nguyenha935

  • fix(providers): classify per-model-quota 403 and…

  • fix(providers): reconcile Kimi K3 vision metadata when synced… (#8250) — @Prudhvivuda

  • fix(providers): the Providers page Learn more button now opens the… (#8284) — @KaynXu

  • fix(api): fold namespace into the flattened Chat tool… (#8295)

  • fix(runtime): sanitize Muse Spark fetch failures before logging and… (#8298) — @backryun

  • fix(sse): Combo middleware preserves OpenAI Responses request… (#8310) — @ridho9

  • fix(api): accept the current compatible-provider… (#8326)

  • fix(api): stop leaking the internal provider UUID in… (#8327)

  • fix(dashboard): show custom provider_nodes providers in… (#8328)

  • fix(api): stop the 2000-token safety buffer from… (#8331)

  • fix(backend): keep combo routing from dispatching image… (#8332)

  • fix(sse): strip third-party-agent signals from the… (#8350)

  • fix(i18n): Restore brand/model proper nouns (Claude, OpenAI,… — @ikelvingo

  • fix(api): estimate inline base64 image tokens instead… (#8368)

  • fix(backend): stop prompt-cache affinity from silently… (#8370)

  • fix(api): accept a missing status query-param on GET… (#8374)

  • fix(resilience): treat an unreachable-proxy… (#8376)

  • fix(backend): make disabling the global per-key proxy… (#8385)

  • fix(dashboard): persist compression engine detail… (#8388)

  • fix(plugins): fire registered+active plugin hooks… (#8395)

  • fix(resilience): cap the connection cooldown after a… (#8396)

  • fix(api): reach synced model_capabilities rows for… (#8429)

  • fix(providers): stop marking a multi-quota-window… (#8431)

  • fix(backend): compute AgentBridge diagnose… (#8466)

  • fix(providers): OpencodeExecutor honors Extra API Keys rotation via… (#8467) — @Prudhvivuda

  • fix(sse): stop combo's "all targets failed" response… (#8486)

  • fix(backend): capability filters fail closed when every combo target… (#8488) — @Prudhvivuda

  • fix(providers): persist a runtime-discovered… (#8491)

  • fix(cli): backup create / backup auto enable —… (#8512)

  • fix(dashboard): Endpoint/API base URL display honors… (#8514) — @rqzbeh

  • fix(client): Absolute fetch("/api/...") and… (#8515) — @rqzbeh

  • fix(cli): create ~/.cache (and set XDG_CACHE_HOME)… (#8519)

  • fix(api): surface a non-blocking warning + startup scan… (#8530)

  • fix(kiro): support profileless Builder ID quota, preserve CLI auth… (#8565) — @nguyenha935

  • fix(test): isolate context-manager unit tests on a… (#8568)

  • fix(quality): check:file-size --update now removes baseline entries… (#8589) — @MumuTW

  • fix(sse): restore Task-Aware Routing config from… (#8604) — @MumuTW

  • fix(sse): task-aware routing defaults use auto/* intents… (#8605) — @MumuTW

  • fix(providers): keep registered OpenCode Go effort aliases in the… (#8610)

  • fix(resilience): report local request-queue expirations as… (#8613)

  • fix(docker): Honor OMNIROUTE_BASE_PATH behind reverse-proxy… (#8615) — @DinonowDev

  • fix(resilience): recover a wedged local request limiter after an… (#8616)

  • fix(providers): Kimi Coding billing-cycle quota errors now recover… (#8632) — @glazec

  • fix(oauth): show GitLab Duo OAuth app / env setup instructions in… (#8688)

  • fix(sse): combo compression-limit resolution falls back to… (#8716) — @DinonowDev

  • fix(db): fall back to Node's built-in SQLite driver when… (#8724) — @epsilonode

  • fix(resilience): keep missing request resources such as Files API ids… (#8756) — @wilsonicdev

  • fix(api): GET /v1/models no longer rebuilds the whole catalog… (#6408 #8833)

  • GHE Copilot: OpenAI-native models now route to <gheUrl>/responses… (#8835) — @hppsc1215

  • CLI: omniroute launch no longer truncates pass-through… (#8837) — @sumanxg

  • Providers: Codex GPT-5.6 Sol/Terra/Luna report the current… (#8838) — @TitoTFP

  • Antigravity: a token refresh now recovers a missing projectId… (#8842) — @HouMinXi

  • CI: check:tracked-artifacts raises the git ls-files… (#8844)

  • OAuth: ANTIGRAVITY_OAUTH_CLIENT_TYPE=web lets a remote… (#8845) — @HouMinXi

  • fix(executors): Vertex AI now routes current-generation Claude models… (#8852) — @wgordon17

  • CLI: omniroute launch-codex no longer corrupts the codex… (#8856) — @sumanxg

  • Providers: AGY model sync now discovers newly announced chat…

  • fix(security): bound JWT-extraction regexes in…

  • Fixed every non-streaming Codex (ChatGPT-account) chat…

  • Logging: a broken stdout/stderr pipe no longer drives a… (#8181)

  • docs: correct three stale references caught by the…

  • Fix the Turbopack next build breaking with "Module…

  • Fix mitm-dnsConfig unit tests failing when the suite… (#6122)

  • muse-spark-web: the WebSocket 401 auth-failure message…

  • Harden the Notion web provider thread-session cache…

  • fix(security): OIDC login gate now requires…

  • Build: the packed tarball boots again — #7191's…

  • Packaging: new check:pack-boot gate packs the real npm tarball,…

  • Packaging: pack-artifact closure tests now cover EVERY npm-shipped…

  • fix(cli): CLI detection now refreshes stale cached results,…

  • fix(docker): Podman deployment guidance now distinguishes local… (#8497)

  • fix(ui): Theme React Flow controls correctly in dark mode,…

  • fix(plugins): the generated plugin host script is now deleted…

  • fix(compression): the Headroom SmartCrusher tabular-compaction guard now… — @SingCJ

  • Skills: register cli-skill-collector in the agent-skills…

  • fix(responses): Make Responses API to Chat Completions downgrades…

  • fix(stream): Responses API clients (/v1/responses) no longer… (#4633)

  • fix(security): bump adm-zip to ^0.6.0 (dev transitive…

  • chore(deps): bump fast-uri≥3.1.3, hono≥4.12.27,…

  • chore(deps): bump dompurify≥3.4.12,…

  • fix(quality): register the two covering unit tests…

  • fix(auth): restore the TICK_MS constant dropped by…

  • fix(guardrails/chat): do not whole-request-reroute…

  • Providers: yuanbao-web no longer forwards a foreign single cookie…

  • Antigravity: the shared token-refresh service now recovers a missing… (#8860) — @HouMinXi

  • Adobe Firefly: gpt-image requests now default detailLevel to… (#8863) — @artickc

  • Combos: deleting a provider connection now clears the combo… (#8865) — @HouMinXi

  • Auto routing: auto/<family> combos (auto/glm, auto/gemini,… (#8866) — @rafaeldrincon

  • fix(oauth): wire Test Connection for xAI OAuth (xai-oauth /… (#8862) — @allanvb

  • Logs: a failed auto/<family> request no longer writes every… (#8867) — @rafaeldrincon

  • fix(usage): correct token/request counting for 30D/90D/YTD/ALL… (#7300) — @growab

  • fix(notion-web): use Chrome TLS impersonation for runInferenceTranscript (#8159) — @artickc

  • fix(bifrost): send v-prefixed transport version, not bare semver (#8194) — @seanford

  • fix(resilience,translator): three release/v3.8.49 base-red regressions + eslint… (#8254)

  • fix(sse): export PROVIDER_BREAKER_FAILURE_STATUSES — fix… (#8258)

  • fix(providers): limit Gemini CLI to legacy OAuth refresh (#8275) — @backryun

  • fix(cli): resolve claude.cmd on Windows in omniroute launch (#8283) — @Dingding-leo

  • fix(resilience): skip terminal connections in token health check sweep (#8286) — @Dingding-leo

  • fix(sanitizer): strip zero-width chars from Anthropic-native streaming… (#8287) — @Dingding-leo

  • fix(i18n): re-sync and complete es-ES translations with latest… (#8289) — @Dragost

  • fix(antigravity): add missing gemini-3.6-flash pricing rows to ag OAuth… (#8290) — @HouMinXi

  • fix(dashboard): stop sidebar RSC prefetch storms (#8292) — @RaviTharuma

  • fix(backend): add structure-aware chat admission (#8296) — @RaviTharuma

  • fix(providers): adapt Kimi nonstream requests internally (#8302) — @RaviTharuma

  • fix(mcp): keep POST SSE responses uncompressed (#8303) — @RaviTharuma

  • fix(api): enforce image generation API key auth (#8306) — @fenix007

  • fix(cpa): isolate credential-pool failures (#8308) — @RaviTharuma

  • fix(sse): suppress by default on Chat Completions (#8309) — @Prudhvivuda

  • fix(translator): cap thinking budget on explicit budget_tokens path (#8312) — @HouMinXi

  • fix(providers): classify HTTP 400 model-unavailable as MODEL_NOT_FOUND… (#8319)

  • fix(providers): carve cookie-auth providers out of terminal 401… (#8321)

  • fix(api): fold namespace into the flattened Chat tool name so… (#8322)

  • fix(providers): route noauth opencode-zen connections through their… (#8324)

  • fix(security): remove quadratic trailing-slash trim in Alibaba… (#8333)

  • fix(memory): self-heal upsertVector/deleteVector from a raced… (#8337) — @hartmark

  • fix(sse): stop stripInternalReasoningPlaceholder from eating… (#8341) — @hartmark

  • fix(compression): rank codex-responses in adaptive-ladder maps (#8381)

  • fix(sse): gate reasoning-placeholder strip to chunks that contain… (#8382)

  • fix(resilience): terminal-skip spares the recoverable GitHub Copilot… (#8389)

  • fix(sse): re-export PROVIDER_BREAKER_FAILURE_STATUSES for the… (#8390)

  • fix(sse): family auto-combos include any backend that serves the… (#8391)

  • fix(sse): cap exact cooldowns only when synthetic — verified… (#8393)

  • fix(devin-cli): update ACP JSON-RPC protocol for Devin CLI 3000.2.x… (#8425) — @MumuTW

  • fix(auth): exclude local CLI providers from tokenHealthCheck… (#8426) — @MumuTW

  • fix(oauth): add devin-cli and agy entries to OAUTH_TEST_CONFIG (#8427) — @MumuTW

  • fix(dashboard): preserve connection health visual on last routed… (#8428) — @MumuTW

  • fix(cursor): bridge native TodoWrite completions (#8432) — @makcimbx

  • fix(guardrails): Vision Bridge describe-fallback ignores unreachable… (#8433) — @guhcostan

  • fix(sse): record tool calls into shared state for… (#8462) — @hartmark

  • fix(oauth): actionable guidance for LAN-origin loopback mismatches… (#8463)

  • fix(hyperagent): sticky thread for agentic tool loops (Claude Code) (#8470) — @artickc

  • fix: combo input-bound, Responses->Chat image strip,… (#8476) — @herjarsa

  • fix(sse): call the real abort-signal helper in the Gemini… (#8485) — @backryun

  • fix(providers): honor Extra API Keys rotation in OpencodeExecutor (#8493) — @Prudhvivuda

  • fix(backend): fail closed when capability filters empty the combo pool (#8494) — @Prudhvivuda

  • fix(providers): resolve native vision for path-shaped multimodal model… (#8495) — @Prudhvivuda

  • fix(hyperagent): default 1M context for fable/opus/sonnet (#8496) — @artickc

  • fix(notion-web): mint fresh thread for new OpenAI sessions with same… (#8511) — @artickc

  • fix(db): classify compressionDetailNormalizers as db-internal in… (#8534) — @MumuTW

  • fix(sse): label HTTP 499 disconnects as client_disconnected (#8552) — @DinonowDev

  • fix(resilience): honor comboCooldownWait for every combo strategy (#8559) — @DinonowDev

  • fix(backend): add error.code to synthOpenAIErrorChunk for guard… (#8570) — @marceli1404

  • fix(video): validate Veo AI Free artifacts before success (#8581) — @brunnolouzada

  • fix(routing): exclude locked-out models from auto-combo candidates (#8586) — @Prudhvivuda

  • fix(api): prevent silent lost updates on concurrent settings… (#8587) — @Prudhvivuda

  • fix(auto-combo): short-circuit expandAutoComboCandidatePool when… (#8598) — @mikeiagents

  • fix(auth): tag internal/loopback-origin failed logins in the audit… (#8606) — @Prudhvivuda

  • fix(providers): declare explicit OpenCode plugin feature-flag defaults (#8608) — @Prudhvivuda

  • fix(ci): use webpack fallback for Node 26 compat build to stop… (#8611) — @Prudhvivuda

  • fix(claude): classify native subscription quota 429 (#8628) — @costaeder

  • fix(windows): request shell when spawning bare qoder binary name on… (#8633) — @Dingding-leo

  • fix(middleware): declare withInjectionGuard's context parameter optional (#8644) — @backryun

  • fix(providers): handle space-separated search queries via… (#8660) — @maxmad64bis

  • fix(dashboard): broken icon allignment in SegmentedControl.tsx (#8679) — @0x2f0

  • fix(providers): deprecate Monster API provider (fixes #8676) (#8691) — @Dingding-leo

  • fix(quality): register the #8494 covering test in stryker… (#8692)

  • fix(sse): clamp max_tokens to the model output cap on every path (#8698)

  • fix(executors): pass TimeoutError reason to controller.abort() in 7… (#8699)

  • fix(api): serve /v1/models stale-first and sanitize its error… (#8703)

  • fix(sse): stop thrashing the provider prompt cache for… (#8705)

  • fix: repair five base-red failures on release/v3.8.49 (#8706)

  • fix(oauth): show GitLab Duo setup before authorize error (#8710) — @DinonowDev

  • fix: hydration mismatch, duplicate React keys, and missing… (#8723) — @Mananz90

  • fix(sse): hide synthetic OpenAI startup reasoning (#8729) — @rinseaid

  • fix(guardrails): check feature flag helper in PIIMasker to honor DB… (#8730) — @Dingding-leo

  • fix(sse): report a stream that completes without any content (#8732) — @backryun

  • fix(plugins): delete the generated host script synchronously so… (#8749) — @MumuTW

  • fix(sse): update ANTHROPIC_PING heartbeat data payload to… (#8762) — @Dingding-leo

  • fix(cloudflare-ai): sync missing free catalog models (fixes #8725) (#8763) — @Dingding-leo

  • fix(bedrock): preserve additionalModelRequestFields in converse… (#8764) — @Dingding-leo

  • fix(docs): use correct kebab-case CLI flags in OpenCode guide… (#8766) — @swingtempo

  • fix(db): repair the extra-migration-dirs test and env contract… (#8773)

  • fix(electron): use NEXT_DIST_DIR when stripping stale native modules (#8794) — @NBN-N3

  • fix(resilience): add UND_ERR_SOCKET to PROXY_UNREACHABLE_ERROR_CODES (#8795) — @Dingding-leo

  • fix(combo): fail open when strict context filter empties… (#8798) — @DinonowDev

  • fix(resilience): honor Cloudflare 1010 retryable:false — skip… (#8800) — @DinonowDev

  • fix(dashboard): restore Usage Model Breakdown column sorting (#8802) — @DinonowDev

  • fix(api): treat cx/* and codex/* as equivalent API-key model… (#8805) — @Prudhvivuda

  • fix(sse): pass real response payload into plugin onResponse hooks (#8806) — @Prudhvivuda

  • fix(types): export web executor types from their module (#8810) — @backryun

  • fix(types): preserve reasoning policy record shapes (#8811) — @backryun

  • fix(types): preserve compression detail config shapes (#8812) — @backryun

  • fix(types): declare virtual chaos combo config (#8815) — @backryun

  • fix(types): type memory skills injection logger (#8816) — @backryun

  • fix(types): align web fallback contracts (#8819) — @backryun

  • fix(types): declare idempotency input contracts (#8820) — @backryun

  • fix(reasoning): sanitize streamed K3 think tags (#8821) — @rinseaid

  • fix(types): normalize Kie transcription results (#8824) — @backryun

  • fix(dashboard): surface quota pool delete failures instead of failing… (#8829)

  • fix(autoCombo): normalize -free suffix in task fitness lookups (#8636) — @Dingding-leo

  • fix(sse): split the reasoning-decision guards so the error arm… (#8641) — @backryun

  • fix(openai-to-claude): skip empty reasoning_content string deltas (#8642) — @Dingding-leo

  • fix(autoCombo): match longest pattern first in lookupStaticFitnessTable (#8651) — @Dingding-leo

  • fix(providers): sync search query to URL query param (#8652) — @Dingding-leo

  • fix(maritalk): refresh supported model catalog (#8663) — @backryun

  • fix(agentSkills): scope command slice before nested subcommands to… (#8667) — @Dingding-leo

  • fix(guardrails): ensure request PII masking respects DB feature flag… (#8715) — @Dingding-leo

  • fix(reasoning): sanitize Kimi Code K3 think tags (#8721) — @rinseaid

  • fix(sse): clamp out-of-range Gemini thinking_budget and learn the… (#8726) — @fuko2935

  • fix(providers): support multi-token search in provider filter (#8731) — @Dingding-leo

  • fix(api): stop JSON.stringify-ing messages before estimateTokens… (#8740)

  • fix(quality): let --update remove baseline entries that already fit… (#8741) — @MumuTW

  • fix: escape angle brackets in i18n messages to prevent… (#8747) — @SteeleHu

  • fix(dashboard): the /home quick-start cards no longer prefetch — #8292…

  • fix(dashboard): Request Logs detail no longer crashes on a structured…

  • fix(dashboard): the logs detail modal stops reopening on first close…

📚 Docs
  • docs(quality): codify retry policy per runner + release-level drift… (#7107)

  • docs(troubleshooting): document Avast/AVG README.md false positive (#7295)

  • docs(perf): add per-endpoint p50/p95/p99 latency + cost budget… (#7336) — @KooshaPari

  • docs: refresh revoked Discord invite + WhatsApp Brasil link (#7604)

  • docs(readme): animated SVG for the 4-tier auto-fallback cascade (#7615)

  • docs: sync provider count to 259 (unblocks docs-counts strict… (#7616)

  • docs(readme): animate pool + combo ASCII blocks as SMIL SVG diagrams (#7626)

  • docs(readme): animate CLI command list + compression flow as SMIL SVGs (#7637)

  • docs(readme): replace free-tier budget mockup with animated SMIL card (#7665)

  • docs(readme): standardize all README tables to full content width (#7666)

  • docs: fix three stale references failing the fabricated-docs… (#7728)

  • docs(readme): unified animated card system — audited v3.8.49 numbers,… (#7769)

  • docs(readme): add Kimi (Moonshot AI) official supporter section (#7770)

  • docs(getting-started): reorder Verify It Works before IDE/CLI setup + add… (#7790) — @swingtempo

  • docs(readme): audit every number against the live code + refresh… (#7795)

  • docs(readme): evolve supporter section into sub2api-style Sponsors… (#7799)

  • docs(readme): contributors 360+ -> 350+ (audited) (#7803)

  • docs(i18n): refresh Polish README and fix relative links (#7807) — @leszek3737

  • docs: add general Web Cookie provider setup guide (#7881) — @arpit-jaiswal-dev

  • docs(guides): document Kaspersky PDM behavioral false positive on the… (#7923)

  • docs: document npm install ERESOLVE/peer/deprecated warnings… (#7988) — @Dingding-leo

  • docs: fix Docker IPv6 connection reset with -p 127.0.0.1 bind… (#7989) — @Dingding-leo

  • docs(readme): Kimi partner tracking links (aff=omniroute) +… (#8028)

  • docs: add AgentRouter multi-provider routing troubleshooting (#8049) — @leninejunior

  • docs(security): correct prompt-injection severity table +… (#8113) — @rafaumeu

  • docs(ops): publish public branching and release model (#8129) — @c4usal

  • docs(i18n): full Russian README rewrite (#8217) — @MonteNegroX

  • docs(readme): re-audit numbers, fix table scroll, refresh contributors (#8243)

  • docs(db): propose pluggable persistence boundary (#8261) — @xiaoyaner0201

  • docs(db): add reproducible SQLite coupling inventory (#8262) — @xiaoyaner0201

  • docs: add public ROADMAP (3.8.5x rail -> 3.9.0 LTS -> 4.0… (#8348)

  • docs: sync env-var contract (chaos panel, notion TLS, grok… (#8362)

  • docs: one golden path across PR template, CONTRIBUTING,… (#8380)

  • docs(claude-code): document unprefixed model IDs and the Ambiguous model… (#8410) — @dvirarad

  • docs(podman): clarify Podman Machine deployment (#8569) — @shixi-li

  • docs(env): document NEXT_PUBLIC_OMNIROUTE_BASE_PATH and… (#8690) — @MumuTW

  • docs(codex): document session affinity and stream idle for long tasks (#8709) — @DinonowDev

  • docs: replace outdated Polish docs with translation from… (#8823) — @leszek3737

  • docs: restore the Polish API_REFERENCE removed by #8823 (#8831)

🧪 Tests & Quality
  • test(build): derive pack-artifact closures for all npm-shipped… (#7081)

  • test(dashboard): dedicated regression guard for #6815 density guarantee (#7291)

  • test(ci): make #6634 selfref guard hermetic — read file from… (#7327)

  • test(ci): mock route bridge surfaces error message, not raw stack (#7354)

  • test(ci): static body in codex e2e mock route bridge (CodeQL #737) (#7558)

  • test(ci): exact-line assert in grok-build config test (CodeQL… (#7628)

  • test(codex): cover image tool output replay (#7704) — @dongwook-chan

  • test(security): exact SAN-entry match in mitm leaf-cert test (CodeQL… (#7824)

  • test: verify keepalive interval cleanup on disconnect,… (#8190) — @rafaumeu

  • test: realign catalog snapshot tests to current deliberate… (#8386)

  • test: hermetic notion thread-session mocks + drop duplicated… (#8392)

  • test(e2e): contract test for the full provider journey (#8444) — @HoneyTyagii

  • test(sse): repair two base-red gates on release/v3.8.49 (#8490) — @backryun

  • test(context): isolate context-manager suite from local DATA_DIR (#8596) — @DinonowDev

🔧 Chores / CI
  • chore(release): gate the sync-back push on release-green --quick (WS0.3) (#7083)

  • chore(ci): gate hygiene — secrets baseline 0, semgrep drop,… (#7099)

  • chore(ops): runner-box janitor + operations runbook (WS3.3) (#7115)

  • chore(ci): promote test:vitest:ui to blocking (suite green after… (#7147)

  • chore(ci): stop dependabot proposing typescript majors —… (#7306)

  • chore(release): script the 0a.0b PR re-home with a verified read-back (#7312)

  • chore(quality): tighten the coverage ratchet to the CI's real numbers (#7326)

  • chore(ci): make the Electron Windows leg advisory with bash stderr… (#7340)

  • chore(deps): bump actions/setup-node from 6 to 7 (#7348)

  • chore(deps): bump codecov/codecov-action (#7350)

  • ci(release-green): add a main-green arm to detect when main goes red (#7355)

  • chore(deps): bump github/codeql-action/analyze from 4.37.0 to 4.37.1 (#7641)

  • chore(deps): bump github/codeql-action/init from 4.37.0 to 4.37.1 (#7642)

  • chore(quality): register #6672 test in stryker tap.testFiles (base-red… (#7652)

  • chore(release): merge-train box-speed suite + --fast mode (#7670)

  • chore: [defer] fix embedded CLIProxyAPI config handling (#6877) — @professional-ALFIE

  • chore: [defer] fix(grok): align responses tool-call shape for… (#6937) — @CitrusIce

  • chore: [needs-vps] feat(dashboard): add per-operator quota row… (#7251)

  • chore: [defer] feat(combo): universal cooldown-aware retry &… (#7301) — @ViFigueiredo

  • chore: Completing Arabic language (#7686) — @mustafa-phd

  • chore: IC2: Cache provider connections by ID + provider nodes (#7744) — @oyi77

  • chore: [Emergency Fix] fix(build): repair release build… (#7772) — @backryun

  • chore: [Part 1/3]refactor(qwen): replace legacy Qwen Code and… (#7866) — @backryun

  • chore: [Part 3/3] feat(qwen): add regional Alibaba and Qwen… (#7882) — @backryun

  • chore: Preserve supported Responses behavior in Chat… (#7894) — @JxnLexn

  • chore: Restore Responses API custom tool calls (#7905) — @JxnLexn

  • chore: Hide internal reasoning replay placeholders (#7912) — @JxnLexn

  • chore(deps): bump js-yaml, brace-expansion, shell-quote, tar… (#7915)

  • chore: i18n(zh-TW): complete Traditional Chinese (Taiwan)… (#8024) — @lunkerchen

  • chore: i18n: bring 40 locales to full parity with en.json (#8031) — @nguyenha935

  • chore(deps): resolve 7 open Dependabot alerts via npm overrides (#8066)

  • chore(deps): resolve 3 more Dependabot alerts (dompurify,… (#8069)

  • chore(dashboard): reframe Kimi partnership as "Open Source Friends" (#8117)

  • chore(quality): fix 2 pre-existing lint/suppression drift issues (#8209) — @hartmark

  • chore: add K3banner-1.png banner asset (#8242)

  • chore(quality): rebaseline accountFallback+combo for #8252 own-growth… (#8252) — @RaviTharuma

  • optimize(chaos+ponytail): i18n ponytail, dedupl dispatch, provider diversity (#8264) — @Moseyuh333

  • chore(deps): bump next to 16.2.11 (9 security advisories) (#8265)

  • chore: Add Alibaba-family media model support (#8266) — @backryun

  • chore: Clicking a provider card hitting back loses scroll… (#8349) — @swingtempo

  • chore: feat(log): Added new visual scrolling log page (#8354) — @hartmark

  • chore: [BUG] enforceOutputTokenBudget ignores combo-resolved… (#8378) — @RCrushMe

  • chore(ci): cancel superseded runs, skip DAST on docs-only PRs,… (#8379)

  • chore(quality): drop stale muse-spark-web allowlist entry + sync… (#8383)

  • chore: i18n(zh-TW): translate missing Reasoning Routing strings (#8423) — @MumuTW

  • chore(deps): bump codecov/codecov-action from 5.5.5 to 7.0.0 (#8452)

  • chore(deps): bump github/codeql-action/analyze from 4.37.1 to 4.37.3 (#8453)

  • chore(deps): bump ossf/scorecard-action from 2.4.3 to 2.4.4 (#8454)

  • chore(deps): bump github/codeql-action/init from 4.37.1 to 4.37.3 (#8455)

  • chore(deps): bump actions/setup-python from 6 to 7 (#8456)

  • chore(sse): drop deprecated baseUrl from open-sse tsconfig for TS… (#8473) — @backryun

  • refactor(sse): resolve open-sse utils/translator type diagnostics for… (#8483) — @backryun

  • refactor(sse): declare the executor execute() result contract (#8489) — @backryun

  • refactor(sse): stop three executors shadowing BaseExecutor.buildHeaders (#8498) — @backryun

  • refactor(sse): narrow three result unions via type predicates (#8499) — @backryun

  • refactor(sse): restore three executor types the runtime already relied… (#8520) — @backryun

  • refactor(dashboard): clear file-size base-red by extracting provider-card… (#8524) — @MumuTW

  • refactor(guardrails): narrow the documented void return at the dispatch site (#8525) — @backryun

  • refactor(sse): type the rerank response adapter's options parameter (#8528) — @backryun

  • refactor(sse): retag the designer-web result unions with string… (#8531) — @backryun

  • refactor(sse): declare the multipart/gRPC frame bodies as… (#8533) — @backryun

  • chore(ci): resync stale no-explicit-any suppression count for… (#8544) — @MumuTW

  • chore(usage): decompose services/usage.ts into per-provider usage/*… (#8545) — @MumuTW

  • chore(validation): decompose providers/validation.ts into validation/*… (#8546) — @MumuTW

  • chore(token-refresh): decompose services/tokenRefresh.ts into tokenRefresh/*… (#8547) — @MumuTW

  • chore(combo): extract pure error predicates and quota status helpers… (#8548) — @MumuTW

  • chore(i18n): normalize zh-TW terminology and sync stale README… (#8554) — @MumuTW

  • refactor(sse): stop isClaudeEventPayload claiming a narrowing it never… (#8557) — @backryun

  • chore(quality): rebaseline file-size for inherited base growth (#8561) — @fenix007

  • chore(deps): patch js-yaml + postcss for 2 high Dependabot alerts (#8572)

  • refactor(tls): update TLSClient instantiation to use… (#8583) — @DinonowDev

  • refactor(sse): peel the bare Response off handleChatCore's union in… (#8647) — @backryun

  • refactor(sse): guard the KIE task-id and callback-url reads at their… (#8661) — @backryun

  • refactor(sse): declare the ArrayBuffer backing on media byte producers (#8665) — @backryun

  • merge: resolve conflicts for #7904 local corpus context… (#8685) — @terrafirmbot-source

  • chore(quality): update baselines after v3.8.49 merge-train (#8686)

  • chore: Dependabot updates (#8695) — @justdoGIT

  • chore: Create AMIT (#8727) — @amitgolan60-coder

  • chore(ci): add release PR build gate (#8735) — @ekinnee

  • chore: Stabilize Notion web sessions and JSON output (#8751) — @0xheycat

  • deps: bump electron from 43.1.1 to 43.2.0 in /electron (#8782)

  • deps: bump the production group across 1 directory with 18… (#8793)

  • refactor(sse): retag the Adobe Firefly IMS token-check union (#8638) — @backryun

  • refactor(db): let the provider-nodes cache keep the row type it… (#8639) — @backryun

  • refactor(sse): use the shared ApiKeyMetadata in reasoningRouting… (#8643) — @backryun

View originalPermalink
How v3.8.49 went

v3.8.48

Added 10
  • Langfuse observability plugin
  • Context requirements config for per-target filtering in combos
  • Icons for 46 providers that were missing images
  • Vendored GCF (Headroom) codec updated to spec v3.2 with nested flattening
  • Shorthand proxy formats and protocol header mode for bulk import
  • OpenVecta AI inference gateway provider
Changed 1
  • Codex bulk-import endpoint now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt) alongside snake_case format
Fixed 9
  • Ship dist/head-response-guard.cjs in the npm tarball to prevent ERR_MODULE_NOT_FOUND crashes on every omniroute boot
  • Fix Electron Windows packaging by spawning npx.cmd through a shell for better-sqlite3 ABI rebuilds
  • Fix Sonar quality gate by ensuring coverage lcov reaches the scanner at coverage/lcov.info
  • Await async isCloudEnabled() gate in Kiro auto-import route so cloud sync respects disabled state
  • Replace dead structuredClone fallback with real JSON fallback in reasoning-split clone
  • Handle async reader.cancel() rejection in codex executor

⚠️ Hotfix release. The published npm package for 3.8.47 crashed on every boot (#7065) and was deprecated — 3.8.48 is the first installable release of the v3.8.47 cycle, so everything listed under [3.8.47] below ships here.

🐛 Bug Fixes
  • fix(build): ship dist/head-response-guard.cjs in the npm tarball — the prepublish prune allowlist lacked it, so every omniroute boot of the published 3.8.47 crashed with ERR_MODULE_NOT_FOUND (3rd occurrence of this class after tls-options/3.8.41); now allowlisted, enforced by check:pack-artifact, and guarded by a closure test that derives every server-ws.mjs sibling import (#7065, #7040)
  • fix(build): Electron Windows packaging — the better-sqlite3 Electron-ABI rebuild now spawns npx.cmd through a shell (Node's CVE-2024-27980 hardening made the shell-less spawn fail with status null on Windows runners, breaking the v3.8.47 desktop build)
  • fix(ci): Sonar quality gate zeroed on new code — the coverage lcov now reaches the scanner at coverage/lcov.info (it read 0% on every scan), the async isCloudEnabled() gate in the Kiro auto-import route is awaited (cloud sync ran even when disabled), the dead structuredClone fallback in the reasoning-split clone is a real JSON fallback, the codex executor handles the async reader.cancel() rejection, deterministic localeCompare sorts, a path-traversal guard in classify-pr-changes.mjs, and the Docker better-sqlite3 rebuild uses npm's bundled node-gyp instead of npx --yes
  • chore(ci): the Sonar quality gate is informational (sonar.qualitygate.wait=false) while the org's SonarCloud plan cannot associate the tuned "OmniRoute way" gate (coverage ≥60 aligned with the repo floor)

📦 Everything from the v3.8.47 cycle ships here

The 3.8.47 npm package was never installable (#7065), so 3.8.48 is the release that actually delivers the whole v3.8.47 cycle — full notes below:

  • 9router Codex import: the Codex bulk-import endpoint (POST /api/oauth/codex/import) now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt + nested providerSpecificData), not just snake_case — normalizeCodexImportRecord maps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged (#6665) — thanks @deadcoder0904. Regression guard: tests/unit/codexBulkImport.test.ts (9router camelCase record, pre-supplied providerSpecificData without an id_token, snake_case-not-overridden, and a full {accounts:[...]} flatten).
✨ New Features
  • feat(plugins): Langfuse observability plugin. (#6577 — thanks @chirag127)

  • feat(combo): context requirements config for per-target filtering in combos. (#6907 — thanks @oyi77)

  • feat(providers): icons for 46 providers that were missing images. (#6926 — thanks @oyi77)

  • feat(compression): vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). (#6838 — thanks @blackwell-systems)

  • feat(proxy): shorthand proxy formats + protocol header mode for bulk import. (#6867 — thanks @growab)

  • feat(provider): OpenVecta AI inference gateway. (#6833 — thanks @hajilok)

  • feat(i18n): Traditional Chinese (zh-TW) localization for frontend and CLI. (#6320 — thanks @lunkerchen)

  • feat(xai): route xAI clients to Grok's native /v1/responses endpoint. (#6709 — thanks @diegosouzapw)

  • feat(routing): per-model web-search/web-fetch interception rules. (#3384, #6814 — thanks @diegosouzapw)

  • feat(release): changelog.d/ fragments — eliminates the CHANGELOG merge-storm cascade. (#6783 — thanks @diegosouzapw)

  • feat(quality): validate-release-green --full-ci reproduces the entire ci.yml static gate set locally. (#6583 — thanks @diegosouzapw)

  • feat(dashboard): sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (src/shared/components/Sidebar.tsx) filters nav sections/groups/items client-side by label as you type, reusing the existing common.search/common.noResults i18n keys (zero new locale edits) and the shared Input icon="search" pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted into filterSidebarSectionsByQuery() (src/shared/utils/sidebarSearch.ts) for isolated unit testing. Regression guard: tests/unit/sidebar-search-filter.test.ts, src/shared/components/Sidebar.search.test.tsx. (#4013 — thanks @crochabe-cyber)

  • feat(combo): auto/* combos gain a strict budget-cap fallback policy — X-OmniRoute-Budget-Fallback: strict (or the persisted config.budgetFallback: "strict") makes an over-budget request fail fast with HTTP 402 instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (cheapest) preserves existing behavior. Builds on the existing X-OmniRoute-Budget/X-OmniRoute-Mode per-request controls (#6023/#6024/#6025), consolidated into resolveRequestAutoControls(). Regression guard: tests/unit/auto-combo-budget-fallback-3470.test.ts. (#3470)

  • Provider/model param filters: config-driven parameter denylist/allowlist per provider/model with auto-learn from upstream 400s (#6649 — thanks @ThongAccount, closes #6625)

  • Per-combo reasoning token buffer toggle: the combo builder now exposes an explicit checkbox for the #3587 reasoning-model max_tokens buffer, defaulting to the existing enabled behavior, so a combo can opt out without hand-editing raw JSON config (#6702 — thanks @xz-dev)

  • feat(dashboard): 9router-parity Routing Strategy settings card on Settings → Routing, plus a per-provider account-routing override on the provider detail page (#6678) — surfaces the existing account round-robin / sticky-limit knobs and adds a new combo-level sticky round-robin (comboStickyRoundRobinLimit, resolved via resolveComboStickyRoundRobinLimit() — per-combo → global combo sticky → account sticky cascade) so combo targets can batch calls per target the same way account fallback already does. A new providerStrategies setting (Zod-validated map, src/shared/validation/settingsSchemas.ts) lets a specific provider override the global fallbackStrategy/stickyRoundRobinLimit without touching the account-wide default, wired into getProviderCredentials() (src/sse/services/auth.ts) ahead of the global fallback. Regression guard: tests/unit/combo-rr-sticky-9router.test.ts, tests/unit/settings-ui-layout-static.test.ts. (thanks @SeaXen)

  • feat(icons): provider logos now resolve local SVG assets first for faster rendering, with a 5-tier fallback chain — local SVG → @lobehub/icons React components → thesvg.org CDN (external SVG for unknown providers) → local PNG → generic AI icon — replacing the previous LobeHub-first order. Adds dozens of first-party provider SVGs and migrates several bitmap logos (continue/copilot/cursor/deepgram/heroku/openclaw/ovhcloud) from PNG to SVG. Regression guard: tests/unit/ui/ProviderIcon-icon-url.test.tsx. (#6317 — thanks @hamsa0x7)

  • Skill Collector CLI detection: new GET /api/skills/collect/detect + POST /api/skills/collect/install (and the cli-skill-collector agent skill) detect which coding CLIs (Claude Code, Codex, Cursor, Copilot, Cline, Hermes, OpenCode, etc.) are installed locally via getCliRuntimeStatus(), match them against GitHub agent-skill repos, and plan an install path per tool — replacing the standalone Skill Collector Python app. Both new routes and GET/POST /api/github-skills now require management auth (requireManagementAuth()) and are loopback-gated (LOCAL_ONLY_API_PREFIXES + SPAWN_CAPABLE_PREFIXES) since the detect route spawns a child process per candidate CLI tool (Hard Rules #15 + #17). The omniroute_github_skills_install MCP tool now reports the honest action: "planned" instead of "installed", matching the REST route (#6294 — thanks @Moseyuh333)

  • ClinePass dual-auth: ClinePass now offers both sign-in methods on its dashboard page — OAuth (reusing the Cline WorkOS flow) as the primary "Connect" path, or a pasted BYOK API key via "Manual API key", instead of only the API-key-only provider shipped in #5942. The registry alias was aligned to cp (matching the OAUTH_PROVIDERS catalog alias) so <alias>/<modelId> routing resolves correctly, the OAuth refresh dispatch now routes clinepass to the shared Cline refresh flow, and the duplicate API-key-only catalog entry was removed to keep ClinePass listed once. Regression guard: tests/unit/clinepass-provider.test.ts. (#6126 — thanks @hajilok)

  • feat(oauth): Kiro/Amazon Q auto-import now supports enterprise External IdP ("Your organization") logins via Microsoft Entra/Okta/Auth0/OneLogin/Ping/Google/Cognito — these org-issued tokens are not AWS SSO tokens (no aorAAAAAG-prefixed refresh token) and can't refresh through the AWS OIDC/Kiro-social path, so tryAwsSsoCache() now detects them (authMethod/provider === "externalidp") and refreshes via the org IdP's own tokenEndpoint (public-client OAuth2 refresh grant, no client secret), persisting TokenType: EXTERNAL_IDP gating so the runtime executor sends the header the AWS CodeWhisperer API requires for these accounts; tokenEndpoint is SSRF-guarded against an HTTPS + known-IdP-host-suffix allowlist. (#6363 — thanks @artickc)

  • Kiro long-lived API key auth: new /api/oauth/kiro/api-key route + KiroService.validateApiKey let a Kiro account be linked with a long-lived AWS CodeWhisperer/Kiro API key instead of the interactive OAuth device flow, with live per-account model discovery (ListAvailableModels, 5-minute cache) layered over the existing static registry fallback (#6587 — thanks @strangersp)

  • Chaos Mode: multi-model parallel/collaborative task execution — dispatches a task to every active provider connection at once (parallel) or chains outputs sequentially so each model builds on the previous one's answer (collaborative), configurable via Dashboard → Chaos Mode (GET/PUT/DELETE /api/chaos/config) and gated per-API-key via a new chaosModeEnabled permission (opt-in — disabled by default globally and per key). POST /api/chaos/run (dashboard session) and POST /api/skills/collect/chaos (external Bearer-token) delegate to a shared executeChaosRun() engine (src/lib/chaos/chaosExecutor.ts) that dispatches in-process via the established synthetic-Request/route-handler pattern (no network hop, no hardcoded port), with a concurrency cap (max 10 parallel), configurable max_tokens (256–128k), a clear error when stream is requested, and collaborative-chain info (provider order + input size). Fixes external Bearer-auth bypass and stale config-cache leakage. Regression guard: tests/unit/chaos-config.test.ts, tests/unit/chaos-executor.test.ts, tests/unit/chaos-api-routes.test.ts. (#6728 — thanks @Moseyuh333)

  • feat(cli): 2 new CLI tool integrations on Dashboard → CLI Tools — omp (Oh My Pi) and letta — each with binary detection, config apply/reset, and a settings card following the existing tool-card pattern. Both settings routes shell out to which omp/which letta to detect the local install, so they're loopback-gated (LOCAL_ONLY_API_PREFIXES, Hard Rules #15/#17) in addition to the shared requireCliToolsAuth() management-auth guard every cli-tools route requires, and route errors through sanitizeErrorMessage(); src/lib/db/omp.ts isolates the omp CLI's own local SQLite reads behind parameterized queries. (Note: the original PR also proposed pi, codewhale, and jcode integrations — those three had already shipped via a separate PR by the time this one was reconciled, so only omp+letta landed here.) Regression guard: tests/unit/db/omp.test.ts, tests/unit/cli-tools-auth-hardening.test.ts, tests/integration/cli-settings-omp.test.ts, tests/integration/cli-settings-letta.test.ts. (#6318 — thanks @hamsa0x7)

  • feat(providers): custom models now support a manual Context Window Override so an operator can correct a provider's misreported context length (e.g. reports 1M when the real limit is 128K) instead of the model silently getting dropped from combo routing once the wrong value lands in the catalog (#4125 — thanks @rucciva). Reuses the existing Feature-5004 model_context_overrides table (source: "manual") — already the priority-0 source getModelContextLimit() (the function combo's context-window filter calls) reads ahead of the models.dev/registry/static catalog — so no new resolver logic was needed, only the missing write path: PUT /api/provider-models now accepts an optional contextWindowOverride (number to set, null to clear), GET surfaces the current value back on each custom-model row, and the provider detail page's custom-model edit form gained a Context Window Override field + badge. Regression guard: tests/unit/provider-models-context-window-override-4125.test.ts.

  • fix(providers): register OpenRouter as a rerank provider so openrouter/cohere/rerank-* models resolve instead of erroring Invalid rerank model (#6574 — thanks @rafpigna)

  • fix(api): HEAD requests no longer hang until client timeout on any route — valid, unknown, authed, or unauthed (#6400), broader follow-up to the route-specific #6517 (/v1/models). Root cause: Next.js 16's App Router route-handler pipeline (next/dist/server/send-response.js) correctly skips piping a Response body for HEAD, but its page-rendering pipeline (next/dist/server/pipe-readable.jspipeToNodeResponse, used for every app-router page/layout render — including the not-found boundary any unmatched path falls through to) has no such check and always streams the full rendered body regardless of method; combined with Node's default keep-alive framing this left some clients unsure whether the (implicitly bodyless) HEAD response had actually finished. A new scripts/dev/head-response-guard.cjs, wired into both the dev/start custom server (scripts/dev/run-next.mjs) and the packaged standalone server (scripts/dev/standalone-server-ws.mjs) at the same tier as the existing http-method-guard.cjs/peer-stamp.mjs wrappers, discards any body bytes written for a HEAD request and forces Connection: close once .end() is called — independent of route existence or auth state, satisfying RFC 9110 §9.3.2. Regression guard: tests/unit/head-request-closes-6400.test.ts.

  • feat(dashboard): Provider Quota page (Dashboard → Quota) fills horizontal whitespace before stacking vertically (#3520) — QuotaCardGrid previously stacked every provider group in a single vertical flex flex-col, and each group's own card grid didn't go multi-column until md (grid-cols-1 md:grid-cols-2 xl:grid-cols-3 2xl:grid-cols-4). Provider groups now flow into a 2-column CSS multi-column layout on very wide (2xl) screens instead of an unconditional vertical stack, and each group's card grid starts at 2 columns immediately (grid-cols-2 md:grid-cols-3 xl:grid-cols-4), reaching higher density sooner on narrower-but-not-mobile viewports. Regression guard: tests/unit/quota-card-grid-horizontal-layout.test.ts (thanks @gdevenyi).

  • feat(ws): the live-dashboard WebSocket server now auto-starts in-process (via instrumentation-node.ts) across every deployment mode — dev, production, Docker, Electron — with no separate sidecar script; the default WS port moved from 20129 to 20132 to avoid colliding with API_PORT in split-port setups, the deprecated OMNIROUTE_DISABLE_LIVE_WS env was consolidated into OMNIROUTE_ENABLE_LIVE_WS (default enabled), and the WS path is now derived from NEXT_PUBLIC_LIVE_WS_PUBLIC_URL's pathname (/live-ws fallback) (#6072 — thanks @ianriizky).

  • feat(compression): new omniglyph engine (context-as-image) — renders system prompt, tool docs, and dense history as compact PNG pages the model reads instead of text (~10× fewer tokens on the converted block; 59–70% end-to-end measured). Works stacked with RTK/Caveman (stackPriority: 90) or standalone (mode: omniglyph); restricted to Claude Fable 5 over the direct Anthropic route, fail-closed gates with skip:<reason> techniques, preview (stable: false, off by default) (#6556). Dependency bumped to omniglyph@^1.0.2 for upstream ReDoS fixes (#6661).

  • feat(sandbox): the skill sandbox gained a container-provider abstraction that auto-detects and uses the best native runtime per host — Apple Container (macOS 26+), WSL container (wslc.exe), OrbStack, Podman — instead of hardcoding docker run, removing the Docker Desktop requirement on macOS/Windows (#6611 — thanks @KooshaPari).

  • fix(sse): skip thinkingConfig for Gemma models on the OpenAI→Gemini path so OpenAI-shape clients no longer get a 400 from Vertex. (thanks @chy1211)

  • feat(xai): route xAI clients to Grok's native /v1/responses endpoint instead of the chat-completions bridge. (thanks @ryanngit)

  • feat(models): add a Settings → AI "Model Overrides" UI plus /api/model-capability-overrides CRUD and a model_capability_overrides table, letting operators set a manual max-output-token override per provider/model (#6727 — thanks @xz-dev).

  • feat(resilience): operator-configurable account rotation policy — a new rotationConfig layer lets operators tune how connections rotate on failure, wired into accountFallback (#6763 — thanks @artickc).

  • chore(cursor): add Grok 4.5 effort/fast model IDs (#6774 — thanks @andrewmunsell).

  • feat(codex): Codex provider model discovery now fetches the live catalog from chatgpt.com/backend-api/codex/models using Codex-shaped headers, falling back to a GitHub-hosted model manifest and then to the local static catalog when the live/GitHub sources are unavailable or return an unexpected shape — new src/app/api/providers/[id]/models/discovery/codex.ts (normalization, version-gating, merge/enrich against the local catalog) covered by tests/unit/provider-models-discovery-split.test.ts and tests/unit/provider-models-route-codex.test.ts (#6776 — thanks @JxnLexn).

  • feat(cursor): register the Opus 4.8, Fable 5, and Sonnet 5 model families for the Cursor Agent provider so the latest Claude/Fable model ids route correctly (#6779 — thanks @andrewmunsell).

  • Changelog fragments (changelog.d/): PRs now add their changelog entry as a new fragment file (changelog.d/{features|fixes|maintenance}/<PR>-<slug>.md) instead of editing CHANGELOG.md — two PRs never touch the same file, structurally eliminating the CHANGELOG-eat merge conflicts that forced a re-sync push + full CI re-run after every sibling merge (O(N²) CI runs in a merge-storm). scripts/release/aggregate-changelog.mjs (npm run changelog:aggregate) folds fragments into the living section at release reconciliation, and check:changelog-integrity now also validates fragment well-formedness. Regression guard: tests/unit/changelog-fragments.test.ts.

  • feat(proxy): add a latency-optimized proxy rotation strategy that ranks pool entries by measured round-trip latency, extending the existing round-robin/random/sticky proxy-pool selection (#6798 — thanks @iamraydoan).

  • feat(fusion): the fusion judge may now draw on its own knowledge and override the panel when every panel answer is wrong or incomplete, instead of being restricted to synthesizing only from panel output (#6804 — thanks @chirag127).

  • feat(dashboard): search box on the Playground's raw model <select> (#4086) — the shared ModelSelectModal (combo builder + CLI-code cards) already had search, but Playground's StudioConfigPane model dropdown stayed a flat unsearchable list, unusable once a provider like OpenRouter contributed 50+ models. Typing now filters the dropdown (Turkish-safe accent/case-insensitive match via matchesSearch), while the currently selected model always stays pinned in the list even if it no longer matches the query, so typing never silently swaps the active selection. Reuses the existing common.search i18n key (already translated in all 42 locales) — no new translation key needed. Regression guard: tests/unit/playground-model-selection-3731.test.ts (filterModelsByQuery), tests/unit/ui/playground-model-search-4086.test.tsx.

  • feat(usage): Antigravity/agy quota widget now surfaces the weekly window alongside the existing per-model 5-hour window (#4017) — the weekly limit isn't part of the per-model retrieveUserQuota response the fetcher already calls; it only appears in a separate, undocumented retrieveUserQuotaSummary RPC that groups models into families ("Gemini Models", "Claude and GPT models") with one weekly bucket per family. A new usage/antigravityWeeklyQuota.ts leaf fetches that RPC (cached, best-effort — a failure or unavailable RPC never breaks the existing per-model quotas) and parses the weekly bucket per group into gemini_weekly/claude_gpt_weekly quota entries, merged into the same quotas map the widget already renders generically. Regression guard: tests/unit/antigravity-weekly-quota-4017.test.ts (bucket parsing, the alternate quotaSummary-nested envelope, and end-to-end merge via getUsageForProvider).

  • feat(codex): Codex CLI compatibility shim — the Responses API response.created/response.in_progress/response.completed payloads now carry a model field (previously absent), and for Codex-CLI-originated requests it echoes the client-requested, effort-suffixed model id (e.g. gpt-5.5-xhigh) instead of the bare upstream id (gpt-5.5), so the Codex CLI status line/model button shows the active reasoning effort (#3697). openaiToOpenAIResponsesResponse (open-sse/translator/response/openai-responses.ts) now threads the upstream model into the Responses event objects; a new isCodexOriginatedHeaders() (open-sse/config/codexIdentity.ts, reusing PR #3481's originator/User-Agent detection) makes chatCore's existing opt-in echoRequestedModelName (#1311) model-echo pipeline fire automatically for Codex clients regardless of the setting, detected by request headers so it still applies when codex/gpt-5.5-xhigh is routed through a combo to a non-codex upstream; echoModelInObject/echoModelInSseLine (open-sse/services/responseModelEcho.ts) now also rewrite the nested response.model field the Responses API uses. /v1/models still returns models: [] for Codex (unchanged). Regression guard: tests/unit/codex-effort-model-echo-3697.test.ts.

  • Z.ai Web (free web-session provider): new zai-web web-cookie provider drives the free chat.z.ai consumer chat UI via a pasted browser session cookie, distinct from the existing API-key zai/glm/glm-cn/glmt providers (api.z.ai) — modeled on the doubao-web/venice-web cookie executors and the pre-existing chatglm-web credential requirement/token-extraction entries. ZaiWebExecutor (open-sse/executors/zai-web.ts) posts to chat.z.ai/api/chat/completions with the cookie forwarded both as Cookie and as Authorization: Bearer <token>, and normalizes both z.ai's internal delta_content/phase SSE envelope and a pass-through OpenAI-shaped choices[].delta frame into standard chat-completion chunks. Registered in WEB_COOKIE_PROVIDERS, WEB_SESSION_CREDENTIAL_REQUIREMENTS, the provider registry (zai-web entry, GLM-4.6/4.5/4.5V models), and tokenExtractionConfig.ts for in-app cookie capture. Regression guard: tests/unit/executor-zai-web.test.ts (16 tests — token extraction, frame parsing for both SSE shapes, streaming and non-streaming aggregation, error paths). (#4056)

  • feat(compression): update the vendored GCF codec behind the Headroom engine to spec v3.2 (nested flattening) (#6837). Homogeneous arrays whose rows carry nested objects/arrays now tabularize via >-prefixed path fields instead of a low-yield per-row fallback, so nested MCP tool-result rows (meta:{...}, tags:[...]) compact like flat rows. On representative shapes the update takes deeply-nested payloads the old codec left near-uncompressed from ~3% to ~32% vs JSON (k8s pods, cl100k_base), with shallow-nested rows seeing a small bump and flat arrays unchanged. Re-vendored from current gcf-typescript (zero runtime deps, MIT, SPDX-marked, generic-profile only); also folds in the [N]: inline-array quoting fix and canonical decimal formatting. Round-trip stays lossless (order-insensitive), and the decoder is hardened against prototype pollution (a __proto__/constructor path segment never mutates Object.prototype, and keys shadowing built-ins like toString now round-trip correctly instead of misparsing). Regression guard: tests/unit/compression/headroom-smartcrusher.test.ts (deep-nested + prototype-pollution cases).

  • feat(providers): Add GPT-5.6 support across OpenAI API, Codex, and ChatGPT Web, including Codex Max/Ultra efforts, VS Code metadata, Fast-tier credit accounting, curated live discovery, the Codex 0.144.1 client identity, and correct chat routing for models that also support image generation (#6862) - thanks @backryun

  • chore(providers): Align emitted Claude Code identity headers, bridge fingerprints, provider profiles, and documented defaults with claude-cli 2.1.207 (#6862) - thanks @backryun

🐛 Bug Fixes
  • fix(dashboard): the Proxy Registry settings page crashed at runtime (ReferenceError: poolLoaded/bulkImportOpen is not defined) — the #6625/#6909 hook-extraction refactors deleted 10 state declarations (poolLoaded, poolSaving, and the 8-member bulk-import family) while ~30 usages remained; all restored (caught by the release E2E; typecheck:core does not cover dashboard TSX — follow-up #7021). (thanks @diegosouzapw)

  • fix(combo): comboStickyRoundRobinLimit now defaults to inherit (null) instead of 1 — the literal default silently shadowed the documented batched round-robin rotation (stickyRoundRobinLimit: 3), flipping every round-robin combo to per-request alternation (#6678 follow-up, caught by the release CI). (thanks @diegosouzapw)

  • fix(ws): the standalone LiveWS startup script exited 0 without ever listening — its bootstrapped child re-spawned with the import-suppressor OMNIROUTE_ENABLE_LIVE_WS=0 and then honored it as an operator disable (#6072 follow-up, caught by the release CI). (thanks @diegosouzapw)

  • fix(api): compression PUT schema accepts every catalog engine. (#6792 — thanks @Pitchfork-and-Torch)

  • fix(compression): surface fallback reasons in the preview response. (#6461, #6519 — thanks @chirag127)

  • fix(providers): fail fast on an empty auto-combo pool instead of a 15s timeout. (#6458, #6546 — thanks @chirag127)

  • fix(compression): honor UI-toggled engines in the stackedPipeline dispatch + surface substitutions. (#6463, #6534 — thanks @chirag127)

  • fix(api): return 400 for missing/invalid messages before model resolution. (#6402, #6515 — thanks @chirag127)

  • fix(providers): enrich the model_cooldown 429 body with a retry_after ISO timestamp + credential count. (#6460, #6523 — thanks @chirag127)

  • fix(sse): default reasoning summary for effort-only Responses requests. (#6807 — thanks @rushsinging)

  • fix(sse): compression no-op treated as zero-savings, not inflation/silent-drop. (#6883 — thanks @chirag127)

  • fix(oauth): Trae OAuth client_id embedded via resolvePublicCred() (Hard Rule #11). (#6870 — thanks @chirag127)

  • fix(sse): combo path no longer trips the whole-provider breaker on a plain 429. (#6868 — thanks @chirag127)

  • fix(api): malformed JSON bodies now return 400 instead of 500. (#6871 — thanks @chirag127)

  • fix(fusion): judge selected from a surviving panel member when no explicit judge is configured. (#6869 — thanks @chirag127)

  • fix(sse): combo model lockout honors the parsed upstream quota reset. (#6863, #6866 — thanks @AgentKiller45)

  • fix(dashboard): logs detail modal no longer reopens on first close. (#6830 — thanks @MikeTuev)

  • fix(usage): honor xAI provider-reported exact cost. (#6711 — thanks @diegosouzapw)

  • fix(kiro): probe IdC region during profileArn discovery, cross-region (recovers #6099). (#6840 — thanks @diegosouzapw)

  • fix(antigravity): sanitize Cloud Code safety settings. (#6839 — thanks @diegosouzapw)

  • fix(translator): defer content_block_start until GLM streams the tool name. (#6730 — thanks @diegosouzapw)

  • fix(translator): strip empty cloud_base_branch from Cursor Subagent tool calls. (#6729 — thanks @diegosouzapw)

  • fix(antigravity): surface aborted Gemini tool calls off end_turn. (#6713 — thanks @diegosouzapw)

  • fix(volcengine): clamp Kimi max_tokens to the Ark endpoint cap. (#6712 — thanks @diegosouzapw)

  • fix(codex): surface capacity errors embedded in 200-OK SSE streams. (#6710 — thanks @diegosouzapw)

  • fix(sse): skip thinkingConfig for gemma models in openai→gemini translation. (#6708 — thanks @diegosouzapw)

  • fix(oauth): avoid bare-email dedup of Codex OAuth logins. (#6706 — thanks @diegosouzapw)

  • fix(sse): unwrap bare {function:{…}} tools in openai→claude translation. (#6704 — thanks @diegosouzapw)

  • fix(db): eliminate a redundant getApiKeyMetadata call in the embeddings route. (#6929 — thanks @oyi77)

  • fix(db): authType filter support in getProviderConnections. (#6946 — thanks @oyi77)

  • fix(i18n): the provider-detail (/dashboard/providers/[id]) visibility + free/paid model filter labels (showVisibleOnly, showHiddenOnly, freeFilterAll, freeFilterFreeOnly, freeFilterPaidOnly, hideAllModels, plus the currently-unused filterVisible/filterHidden/filterByVisibility) rendered as the literal __MISSING__:<english> sentinel in 15 locales, including pt-BR (#6694) — providerText() (providerPageHelpers.ts) checks t.has(key) before falling back to clean English, and t.has() returns true even when the stored value is the __MISSING__: sentinel scripts/i18n/sync-ui-keys.mjs writes when mirroring keys across locales, so the sentinel rendered verbatim instead of the fallback. Disjoint key set from #6290 (filterAll/filterActive/filterError/filterBanned/filterCreditsExhausted). All 9 keys now carry real translations across the 15 affected locale files (it, ja, ko, mr, ms, nl, no, phi, pl, pt, pt-BR, ro, ru, sk, sv). Regression guard: tests/unit/i18n-provider-visibility-filter-keys-6694.test.ts.

  • fix(cli): the dashboard's Claude Code CLI card could report "Not detected"/"Not installed" even when Claude Code was genuinely installed and previously used (#6701) — getCliRuntimeStatus() (src/shared/services/cliRuntime.ts) determined installed purely from binary resolution (known install paths + a where/which PATH search), with no fallback when that lookup fails for reasons unrelated to whether the CLI is actually installed (stale PATH inherited by a long-running/background process, the binary having moved, an install method not yet catalogued, etc.) — even though ~/.claude/settings.json on disk proves the tool was installed and used before. Upstream 9router's equivalent route already has this exact fallback. A new withSettingsFallback() (src/shared/services/cliInstallFallback.ts) restores 9router parity: when the binary lookup's own reason is "not_found" (never for deliberate security rejections like unsafe/relative env overrides or symlink escapes) and the tool's settings file exists on disk, installed now reports true. Regression guard: tests/unit/repro-6701-claude-detect-fallback.test.ts.

  • fix(cli): per-agent AgentBridge DNS toggle was broken for 8 of the 9 supported agents, and a failed MITM startup step could orphan the spawned proxy child — addDNSEntry/removeDNSEntry (src/mitm/dns/dnsConfig.ts) always resolved the legacy Antigravity default hosts regardless of which agent's toggle was flipped, so enabling DNS for Cursor/Codex/Claude Code/etc. silently added only daily-cloudcode-pa.googleapis.com while the DB recorded dns_enabled=true for the selected agent. Both functions now accept an optional agentId and resolve hosts via ALL_TARGETS; POST /api/tools/agent-bridge/agents/[id]/dns passes the route's id through and now returns 404 for an id that doesn't match a known target instead of silently falling back. Separately, startMitmInternal() (src/mitm/manager.ts) now wraps generateCert() (log + rethrow), the provisionDnsEntries() call, and the PID-file write in try/catch so a mid-startup failure can't orphan the already-spawned MITM child process. On Windows, addDNSEntries/removeDNSEntries also batch every missing/present entry into a single elevated PowerShell invocation instead of one UAC prompt per host line. Regression guard: tests/unit/dns-config-generic.test.ts (agent-specific resolution + batching), tests/unit/agent-bridge-dns-route-validation.test.ts (404 for unknown agent id). (#6338 — thanks @hamsa0x7)

  • fix(guardrails): Vision Bridge's individual-model auto-reroute (route an image-bearing request straight to a vision-capable model instead of describe-then-forward) could bypass a policy-restricted API key's model allowlist/budget (#6640) — VisionBridgeGuardrail.preCall() (src/lib/guardrails/visionBridge.ts) swaps body.model to the best available vision-capable model, but that swap happens in the guardrail pipeline AFTER chat.ts already called enforceApiKeyPolicy() against the ORIGINAL model, so a key scoped to a narrow allowedModels list could still execute against an unvetted (and possibly costlier) vision model the reroute picked. chat.ts now re-validates any guardrail-driven model change against the same per-key allowlist (isModelAllowedForKey) before honoring it, falling back to the original already-approved model when the reroute target is not allowed. The reroute path also now honors an explicit settings.visionBridgeModel operator override (previously ignored, unlike the combo/describe path a few lines below it, which already respects it via getVisionBridgeConfig). Regression guard: tests/unit/guardrails/visionBridge.test.ts (22 tests). (thanks @herjarsa)

  • fix(auth): an API key restricted via allowedModels/allowedCombos could bypass that restriction entirely over the Codex Responses-over-WebSocket bridge (#6564) — prepare() in src/app/api/internal/codex-responses-ws/route.ts authenticated the WS bridge's API key (authenticate()/authorizeWebSocketHandshake()) and honored allowedConnections, but never called enforceApiKeyPolicy(), the same model/combo policy gate the HTTP /v1/responses path enforces via handleChat() — so a key scoped to e.g. combo/model-1.0 could still reach a direct Codex model like gpt-5.5 through this transport, as long as an eligible Codex OAuth connection existed. The bridge's WS auth token arrives via query params (api_key/token/access_token), not a normal Authorization header, so a new enforceCodexWsApiKeyPolicy() builds an equivalent Request carrying an explicit Authorization: Bearer <apiKey> header and calls enforceApiKeyPolicy() against the CLIENT-requested model, before any Codex-specific model remapping or credential selection. Regression guard: tests/unit/codex-ws-policy-enforcement-6564.test.ts (a model-restricted key is rejected 403 before reaching credential selection; a combo-restricted key is rejected 403 requesting a disallowed combo; a key that DOES allow the requested model still proceeds past policy). (thanks @Squawk7777 for the report and an independent fix via #6565)

  • fix(security): loopback-gate /api/middleware/* so a leaked JWT over a tunnel can't install or trigger a middleware hook — middleware hooks compile + run arbitrary JS via new vm.Script on the request hot path (src/lib/middleware/registry.ts), the same RCE class as the already-gated /api/plugins/*; /api/middleware/ is now in LOCAL_ONLY_API_PREFIXES so loopback enforcement runs unconditionally before any auth check (Hard Rules #15 + #17). Regression guard: tests/unit/route-guard-middleware-local-only.test.ts. (#6541) — see PR. (thanks @developerjillur)

  • fix(startup): AgentBridge's MITM server no longer fails to start with ROUTER_API_KEY is required on a normal install (#6403) — POST /api/tools/agent-bridge/server resolved the spawned MITM child's router key from only an explicit apiKey body field (never sent by the AgentBridge UI — the schema has no such field) and the ROUTER_API_KEY env var (unset by default), so startMitm() always received "" and the child hard-exited, even though OmniRoute already had a usable API key in its own DB. A new resolveRouterApiKey() now falls back to pickApiKeyForInternalUse() (the same DB-backed selector the combo-health-check / cloud-sync internal probes use), resolving in order: explicit key → ROUTER_API_KEY env → an existing DB key. Regression guard: tests/unit/agentbridge-mitm-router-key-6403.test.ts.

  • fix(providers): deploying a Cloudflare relay Worker from Dashboard → System → Proxy pool → Cloudflare relay failed immediately with Cloudflare Worker upload failed: Content-Type must be one of: application/javascript, text/javascript, multipart/form-data, even with a valid token/account (#6416) — the Worker-script upload built a native FormData and let fetch derive the multipart Content-Type automatically, but in production globalThis.fetch is patched with node_modules/undici's own fetch (open-sse/utils/proxyFetch.ts), whose FormData/Request classes differ from the runtime's global FormData (same cross-realm class mismatch already fixed once for image edits in #3273); passing a native FormData instance through undici's patched fetch made it serialize the body as the literal string "[object FormData]" with Content-Type: text/plain;charset=UTF-8, which Cloudflare rejects outright. buildCloudflareWorkerUploadRequest() (src/lib/proxyRelay/cloudflareWorkerScript.ts) now builds the multipart body as a raw Buffer with an explicit boundary and Content-Type: multipart/form-data; boundary=… header, accepted verbatim by any fetch implementation. Regression guard: tests/unit/cloudflare-worker-upload-content-type-6416.test.ts + updated tests/unit/relay-deploy-5128.test.ts.

  • fix(security): SSRF-guard the provider-validation probes so they can no longer be used as an open relay to cloud-metadata endpoints — directHttpsRequest() (web-cookie / NVIDIA / Z.AI validation, all with a caller-controllable baseUrl) ran with guard:"none" + allowRedirect:true; it now applies getProviderValidationGuard() (default block-metadata: LAN/localhost allowed, 169.254.169.254/link-local IMDS rejected, opt-out via OMNIROUTE_ALLOW_PRIVATE_PROVIDER_URLS) and allowRedirect:false so a provider can't 3xx-redirect the probe to metadata past the initial-URL guard. Regression guard: tests/unit/provider-validation-ssrf-guard.test.ts. (#6542) — see PR. (thanks @developerjillur)

  • fix(startup): AgentBridge's MITM proxy served a mismatched cert for 3 of the 4 antigravity/cloudcode-pa hosts it terminates TLS for, breaking interception (#6494) — src/mitm/server.cjs's TARGET_HOSTS decrypts all 4 hosts locally (daily-cloudcode-pa.googleapis.com, cloudcode-pa.googleapis.com, daily-cloudcode-pa.sandbox.googleapis.com, autopush-cloudcode-pa.sandbox.googleapis.com), but src/mitm/cert/generate.ts's self-signed cert only carried a SAN entry for the first host — a request to any of the other 3 got served a cert whose CN/SAN didn't match (confirmed via curl -k https://cloudcode-pa.googleapis.com/ showing CN=daily-cloudcode-pa.googleapis.com). generateCert() now sources its host list from ANTIGRAVITY_TARGET.hosts (src/mitm/targets/antigravity.ts, the single authoritative registry already kept in lock-step with server.cjs/dnsConfig.ts/mitmToolHosts.ts by their own drift tests) and emits a SAN entry for all 4 hosts instead of hard-coding a second, incomplete copy. Regression guard: tests/unit/agentbridge-antigravity-cert-hosts-6494.test.ts (asserts the host list covers all 4 hosts and that the real generated cert's SAN includes each one).

  • fix(resilience): a priority combo never fell back when a target masked credit/quota exhaustion behind an HTTP 200 (#6427) — validateResponseQuality() (open-sse/services/combo/validateQuality.ts) only inspected the response body's top-level error field when choices was ALSO missing/empty (the narrower #3424 empty-completion case); a masked 200 that echoed a non-empty stub choices alongside a structured error object, or a known exhaustion phrase (e.g. "insufficient credits", "quota exceeded") in the error envelope, slipped through as "valid" and the combo kept returning the dead target's response forever instead of failing over. The quality check now inspects the error envelope — a top-level OpenAI-shape error object, or a bounded, case-insensitive exhaustion-phrase match against error.message/error.code/error.type/top-level message/detail — unconditionally, before any shape-specific branch, and regardless of whether choices/output also look structurally present. The check never inspects choices[].message.content, so a legitimate completion that merely mentions "quota" or "credits" in assistant prose is not misclassified. Regression guard: tests/unit/masked-200-exhaustion-fallback-6427.test.ts.

  • fix(security): fail-closed CORS for the cookie/session-authed cloud-agent management routes — getCloudAgentCorsHeaders() reflected any caller's Origin and paired it with Allow-Credentials: true (a CSRF/exfil hole); it now defers to the central allowlist (resolveAllowedOrigin), echoes only an allowlisted origin with Vary: Origin, and emits Allow-Credentials only for an explicitly allowlisted origin — never for a CORS_ALLOW_ALL wildcard echo. Regression guard: tests/unit/cloud-agent-cors-failclosed.test.ts. (#6543) — see PR. (thanks @developerjillur)

  • fix(compression): adaptive-compression ladder ranked 6 real catalog engines (ccr, ionizer, relevance, llmlingua, llm, read-lifecycle) as if they didn't exist (#6533) — ladder.ts's AGGRESSIVENESS and REDUCTION_FACTOR maps only covered the 7 engines wired into DEFAULT_LADDER (session-dedup/rtk/headroom/lite/caveman/aggressive/ultra); every other engine registered in open-sse/services/compression/engines/index.ts — including ccr/llmlingua, which the ladder doc comment already says are intentionally addable via ladderOverride — fell through to aggressivenessOf()'s ?? 0 default (same rank as "off") and expectedReductionFactor()'s generic ?? 0.9 fallback, so floor-mode escalation could not rank or escalate past them once added to a custom ladder. Both maps now carry entries for all 6 missing real engines, rescaled ×10 (off:0ultra:70) and placed by each engine's documented stackPriority (ionizer between rtk/headroom, relevance before caveman, llmlingua/llm between aggressive and ultra, etc.); mcpAccessibility — named in the report — is not a registered CompressionEngine (it's a separate MCP tool-response truncation mechanism) and was correctly left out. Regression guard: tests/unit/ladder-engine-maps-6533.test.ts (asserts every id from listCompressionEngines() ranks above "off" with a non-default reduction factor). (thanks @chirag127)

  • fix(api): tool-call arguments could render as [object Object] sequences instead of the real JSON through the /anthropic (Anthropic-shape /messages) routing path (#6459) — appendToolCallArgumentDelta() (open-sse/utils/toolCallArguments.ts), the shared accumulator the streaming openai-to-claude response translator, openai-responses translator, and responsesTransformer all call to build up a tool call's arguments/input_json_delta buffer, treated any non-string incoming fragment as an empty string. Some upstreams deliver the full tool_calls[].function.arguments value as an already-parsed JSON object/array instead of the OpenAI-contracted JSON-encoded string; the old code silently discarded that fragment, leaving tool_use.input empty, and left downstream buffers open to a plain string coercion of the object ([object Object]) once client-side concatenation kicked in. appendToolCallArgumentDelta() now JSON.stringify()s a non-string, non-null object/array fragment into a valid JSON fragment instead of dropping it, so the assembled partial_json always parses back into the original structured value. Regression guard: tests/unit/anthropic-toolcall-args-6459.test.ts. (thanks @chirag127)

  • fix(providers): fusion combo returned the opaque "All fusion panel models failed" 503 even when only a minority of panel members were actually cooling down / rate-limited, and a user-supplied fusionTuning.minPanel=1 was silently overridden (#6454) — handleFusionChat() hard-clamped the quorum floor via Math.min(Math.max(2, cfg.minPanel), panel.length), so an operator-configured minPanel=1 never took effect: collectPanel()'s straggler-grace timer only starts once ok >= minPanel, and with the floor forced to 2 a single fast success plus N slow-failing stragglers never reached quorum, so the panel sat waiting instead of degrading to the survivor. Per-member failure reasons (straggler_dropped/timeout/threw/status_XXX/empty_content/unparseable) were also logged server-side but never surfaced in the 503 body, leaving operators unable to tell a rate-limit fan-fail from a broader outage. Fixed by honoring Math.max(1, cfg.minPanel) and threading a failures: Array<{ model, reason }> collector into the 503 message (model=reason per entry) — production fix already merged via #6521; this entry backfills the missing CHANGELOG bullet and adds an 11-member, fusion-free-scale regression test matching the original repro shape (a cooling minority must not sink a healthy majority; a genuinely all-failed panel still returns the documented 503). Regression guard: tests/unit/services/fusion-min-panel-and-failure-detail.test.ts + tests/unit/fusion-partial-panel-failure-6454.test.ts. (thanks @chirag127)

  • fix(providers): fusion combo strategy silently returned a panel member's raw answer instead of the configured config.judgeModel synthesis (#6455) — handleFusionChat()'s single-survivor "degrade gracefully" path (added for #6454) returned the lone panel answer directly whenever only one panelist succeeded, regardless of whether an explicit judgeModel was configured; with the default minPanel: 2 and a 2-model panel, any single flaky/rate-limited panelist forced this path on every request, so the configured judge (e.g. auto/claude-opus) was never invoked and the client-visible .model reflected whichever panelist happened to survive. The judge is now still invoked to synthesize a lone surviving answer whenever judgeModel is explicitly configured; the cheap direct-answer shortcut is kept only for the implicit case (no judgeModel set, where the "judge" is just panel[0]). Regression guard: tests/unit/fusion-judge-model-6455.test.ts + updated tests/unit/combo-fusion-strategy.test.ts. (thanks @chirag127)

  • feat(combo): sanitized diagnostic trace on an auto-combo terminal failure — instead of an opaque 503, a terminal combo failure now returns a whitelist-projected trace (candidate pool size, attempted count, excluded provider/reason codes, attempt order, and a terminal-reason code) via the new errorResponseWithComboDiagnostics()/sanitizeComboDiagnostics() in open-sse/utils/error.ts — provider/model ids and enumerated reason codes only, never keys/tokens/bodies, length- and count-capped. A reasoning-budget-exhausted panel now returns an actionable "increase max_tokens" message rather than a blind retry-limit 503. Regression guard: tests/unit/combo-diagnostics-trace.test.ts. (#6545) — see PR. (thanks @developerjillur)

  • fix(providers): image/diffusion models discovered from an upstream catalog (e.g. HuggingFace's live /v1/models) are no longer advertised as chat models (#6457) — the chat catalog builder defaulted synced models with no modality info to endpoints: ["chat"], so huggingface/stabilityai/stable-diffusion-xl-base-1.0 showed up in the chat /v1/models listing and returned 400 "not a chat model" when called. catalog.ts now skips any synced model already registered as an image model for that provider (via the new isRegisteredImageModel()), leaving getAllImageModels() to list it with the correct type: "image". Regression guard: tests/unit/image-model-not-in-chat-catalog-6457.test.ts.

  • fix(resilience): combo session stickiness never released a pin on a credits-exhausted/banned/expired account, permanently defeating failover for that conversation (#6692) — applySessionStickiness() (open-sse/services/combo/sessionStickiness.ts) gated the sticky pin only on 5h/weekly usage-percentage headroom, which is orthogonal to account availability: a credits_exhausted/banned/expired connection, or one still inside its rateLimitedUntil cooldown, reports perfectly healthy headroom, so the pin was force-promoted back to the front of the target list on every subsequent turn. clearStickyBinding() also had zero call sites in combo.ts's failure paths, so a quality-validation-rejected 200 (a masked daily-cap refusal) never released the pin either. The gate now also resolves the bound connection's terminal status/cooldown via a new injectable fetcher seam (fail-open on lookup errors, mirroring the existing saturation fetcher), and combo.ts's two dispatchers (handleComboChat/handleRoundRobinCombo) release the pin immediately at both their connection-exhaustion classification point and their quality-validation-failure branch via the new releaseStickyPinOnFailure(). Regression guard: tests/unit/repro-6692-sticky-terminal.test.ts + extended tests/unit/combo-session-stickiness.test.ts.

  • fix(test): replace the bare expect(true).toBe(true) tautology in playground-api-tab.test.tsx's SSE test and close the check:test-masking gap that let it slip through for a full cycle (#6404) — a prior pass (#6548) had already swapped the literal to expect(sendBtn).toBeDefined(), but that stayed just as vacuous: the test's fetch mock returned an empty /v1/models list, so ApiTab's Send button is always disabled (!selectedModel) and the SSE branch never runs — the "SSE infra is verified" comment was never true. The test now mocks a real model, drives the model <select> to enable Send, asserts sendBtn.disabled === false before clicking, and asserts the streamed SSE delta ("Hello!") actually reached the response editor. Root cause on the detector side: check-test-masking.mjs's tautology subcheck only compares base-vs-HEAD counts within a PR's own diff (headExtTaut > baseExtTaut) and no-ops locally when GITHUB_BASE_SHA/GITHUB_BASE_REF are unset ("sem base ref — pulando") — so a tautology merged once, or checked with a bare local run, was invisible forever after. Added a new always-on, PR-independent absolute-floor scan (scanBareTautologies + countBareTautologies) over every git-tracked test file for the bare expect(true).toBe(true) / assert.equal(1,1) / assert.strictEqual(1,1) patterns specifically (deliberately excluding assert.ok(true), which has ~15 pre-existing verified-legitimate try/catch-fallback uses repo-wide and stays governed by the lenient diff-only subcheck) — verified zero pre-existing hits repo-wide once this file was fixed, so the new floor is safe to enforce unconditionally. Regression guard: tests/unit/check-test-masking.test.ts (new scanBareTautologies/countBareTautologies cases) + tests/unit/ui/playground-api-tab.test.tsx. (thanks @chirag127)

  • fix(oauth): Codex/ChatGPT (and every other OAuth provider) connection stays stuck showing "Auth Failed" even after a genuinely successful token refresh (#6352) — updateProviderCredentials() (the shared onPersist callback for the manual "Refresh token" route, the reactive per-request refresh in chat.ts, and the Codex/Claude auth-file importers) correctly reused the stored refresh_token, persisted the new access_token, and replaced a rotated refresh_token, but never cleared the stale testStatus/lastError*/errorCode fields left over from a prior expired/invalid refresh or upstream 401/403 — only the separate background health-check sweep did that clearing. A successful refresh now resets testStatus to "active" and clears lastError, lastErrorAt, lastErrorType, lastErrorSource, and errorCode (an explicit testStatus from the caller still wins). Regression guard: tests/unit/codex-oauth-refresh-persist-6352.test.ts.

  • perf(health): short-TTL (1s) cache for the frequently-polled GET /api/monitoring/health payload — rebuilding it every request (DB reads + status aggregation across 8 subsystems) was wasteful under rapid polling; the cache stays near-real-time and is invalidated immediately on DELETE (circuit-breaker reset) so a manual reset is reflected at once. Regression guard: tests/integration/monitoring-health-cache.test.ts. (#6553) — see PR. (thanks @developerjillur)

  • fix(resilience): headroom combo routing did not always select the Codex account with the most free quota (#6379) — orderTargetsByHeadroom (open-sse/services/combo/quotaStrategies.ts) already loaded the per-connection DB snapshot (with decrypted credentials) via expandTargetsByQuotaAwareConnections, but discarded it before calling getSaturation; for Codex, fetchCodexSaturation forwards straight to fetchCodexQuota(connectionId, connection), which needs connection (or a prior registerCodexConnection() call, which never happens before headroom ranking runs) to read accessToken — so it returned null for every candidate, saturation failed open to 0 across the board, and ranking fell back to the original combo order regardless of actual free quota. getSaturation() and the headroom SaturationFetcher seam now accept and thread the loaded connection snapshot through to fetchCodexQuota. Kilo's dup flag vs #5903 was a false positive — that issue is about session-sticky reset-aware/least-used selection, not headroom's Codex saturation lookup. Regression guard: tests/unit/headroom-codex-quota-snapshot-6379.test.ts. (thanks @eidoog)

  • fix(providers): custom models a provider actually has are no longer dropped from the Free Provider Rankings when both "Configured only" and "Available only" filters are applied (#6368), follow-up to #6150 — freeProviderRankings.ts::getProviderModels() only ever walked the static open-sse/config/providerRegistry.ts catalog, so a user-added custom model (e.g. a Puter claude-fable-5 model saved as "Claude Fable 5") never entered the candidate model list the ranking scores against, and could never survive the #6150 configured/available filters even when actually configured and available. It now additively merges the provider's custom models (db/models.ts::getCustomModels) into that candidate list via a new pure, de-duping mergeProviderModels() helper, before scoring/filtering runs — catalog free/paid filtering elsewhere is untouched. Regression guard: tests/unit/free-provider-rankings-custom-models-6368.test.ts. (thanks @shabeer)

  • fix(playground): accept a valid dashboard session for GET/POST /api/playground/presets under REQUIRE_API_KEY=true — the Playground page calls this route with a cookie/session and no API key, which previously 401'd the authenticated dashboard; checkAuth() now accepts a management/dashboard session (requireManagementAuth) as an alternative to an API key, while a presented API key must still be valid and the anonymous-allowed default is preserved. Regression guard: tests/integration/presets-dashboard-auth.test.ts. (#6554) — see PR. (thanks @developerjillur)

  • fix(providers): cloudflare-ai no longer silently drops image/non-text content parts (#6390) — transformRequest()'s flattenContent() (added for #2539 to satisfy the Workers AI /ai/v1/chat/completions plain-string content requirement) mapped any non-text OpenAI content part (e.g. image_url) to "" and joined the rest, so a request carrying an image quietly went out as text-only with the attachment gone and no error surfaced. It now throws a clear error on the first non-text part instead of dropping it silently, which the existing top-level chatCore.ts catch already routes through buildErrorBody()/sanitizeErrorMessage() (same pattern as buildUrl()'s missing-Account-ID error). Regression guard: tests/unit/cloudflare-ai-image-parts-6390.test.ts.

  • fix(providers): stop Antigravity connections from falsely reporting all-accounts quota-exhausted (#6295) — genericQuotaFetcher.ts::percentUsedForQuota() ignored the fractionReported flag and defaulted an unreported model's remainingPercentage to 0, which computed as 100% used; since convertUsageToQuotaInfo() takes the worst-case window across a connection, a single model with no reported fraction dragged the whole account into limitReached and quotaPreflight skipped it. percentUsedForQuota() now returns null (unknown, window ignored) whenever fractionReported === false, before falling back to remainingPercentage. Regression guard: tests/unit/generic-quota-fetcher.test.ts.

  • fix(fusion): the fusion judge no longer replays a panel member's answer via an idempotency-key collision — fusion's panel + judge sub-requests re-enter chatCore sharing the client's headers, so they derived the same Idempotency-Key/x-request-id and a panel answer saved under the key was replayed by the judge's check ~1ms later (inside the 5s window), returning a panel member's answer instead of the judge synthesis (observed live on nexa/conversation-fusion). composeIdempotencyKey() now namespaces the key by target provider/model + a digest of the request messages, so sub-requests can't collide while a genuine client retry (same key/model/body) still replays. Regression guard: tests/unit/idempotency-fusion-collision.test.ts. (#6558) — see PR. (thanks @developerjillur)

  • fix(providers): grok-cli (Grok Build) now strips reasoning_effort/reasoning before forwarding the request (#6288) — Claude Code sends reasoning_effort on every request (routing the Opus slot), which Grok Build's upstream chat-proxy endpoint rejects with a 400; transformRequest()'s existing UNSUPPORTED sampling-param strip list (#5273) never covered it. Regression guard: tests/unit/grok-cli-reasoning-strip-6288.test.ts.

  • fix(cli): omniroute serve no longer hangs silently on a readiness timeout (#6321) — the child server's stdout was piped to "ignore" whenever --log/OMNIROUTE_SHOW_LOG wasn't set (the default), discarding any debug output, and runWithSupervisor's waitForServer(...).then((up) => { if (up) {...} }) had no else branch, so a boot that never became ready produced zero further output after "⏳ Starting server...". Stdout is now buffered alongside stderr (ServerSupervisor.getRecentLog()), and a timeout prints a clear diagnostic plus the buffered output instead of staying silent. Does not by itself explain why boot never completes on a given machine — see the issue for further reproduction. Regression guard: tests/unit/cli-serve-readiness-timeout-6321.test.ts.

  • fix(pricing): Pricing Sync dashboard no longer stuck on "Next Sync: Never" / "Synced Models: 0" (#6325) — pricingSync.ts kept sync state (lastSyncTime, lastSyncModelCount) in module-level vars, but the background periodic sync (instrumentation-node.ts) and the dashboard status route (/api/pricing/sync) each import the module from separate Next.js standalone webpack chunks, giving each its own independent state; getSyncStatus() read the (empty) API-route instance's vars. Sync status is now additionally persisted to a new pricing_sync_status key_value namespace and getSyncStatus() falls back to it when the local module instance never ran a sync itself. Regression guard: tests/unit/pricing-sync-cross-instance.test.ts.

  • fix(api): stop spuriously 403-ing "Invalid request origin" on POST /api/providers/health-autopilot/actions for Docker/LAN dashboard requests (#6277) — the route carried a duplicate per-route validateBrowserMutationOrigin check re-added by the v3.8.42 release squash after PR #5278 centralized origin enforcement in the authz pipeline; the pipeline strips PEER_IP_HEADER before forwarding, so the stale duplicate check could no longer resolve the LAN "direct-local-host" candidate and rejected legitimate same-origin LAN mutations (e.g. clicking "remove cooldown" when accessed via a LAN IP). Removed the duplicate check — origin validation is now solely enforced by the centralized pipeline check, which already handles this case correctly. Regression guard: tests/unit/serial/provider-health-autopilot.test.ts.

  • fix(resilience): a bare, unrecognized 403 from a no-credential (authType:"none") provider like mimocode or theoldllm no longer permanently bans the connection (#6315, #6345) — classifyProviderError()'s 403 branch only exempted apikey providers from the terminal FORBIDDEN classification, so these free/stateless proxies (no real account/credential to revoke) fell through to FORBIDDEN on the first unmatched 403 and got isActive:false, testStatus:"banned" with no cooldown or retry. The exemption now also covers authType:"none" providers, returning null (recoverable) so the existing connection-cooldown/retry layer handles it. Regression guard: tests/unit/errorClassifier-noauth-403-6315.test.ts.

  • fix(providers): the Auggie (Augment CLI) executor no longer fails on Windows with spawn EINVAL (#6304) — the global-npm install exposes auggie as a .cmd shim, which Node's child_process.spawn cannot launch on win32 without shell: true. Both spawn sites (streaming + the auggie --version test) now go through a shared buildAuggieSpawnOptions() that sets shell: process.platform === "win32"; the argv (built by buildAuggieArgs() with a registry-validated model and a trailing -- end-of-options marker) is unchanged, so the argument-injection surface stays closed on non-Windows. Regression guard: tests/unit/auggie-win32-spawn-6304.test.ts.

  • fix(api): the dashboard "Test model" action is now a clean connection test (#6240) — modelTestRunner sent its probe request without an explicit compression override, so whenever the operator's global compression.enabled flag was on the test call inherited compression (and any Output-Styles system prompt), polluting the result. The internal test requests now send X-OmniRoute-Compression: off, and chatCore honors an explicit off header even when compression.enabled is globally true. Regression guards: tests/unit/model-test-runner-compression-off-6240.test.ts, tests/integration/test-model-compression-off-6240.test.ts.

  • fix(startup): an update/restart could crash the whole server at boot with TypeError: Cannot create property 'message' on string 'Database closed', masking the real failure and 500-ing every request until manually restarted (#6560, plausibly the root cause of #6594's post-upgrade 500) — driverFactory.ts::preInitSqlJs() cached its sql.js WASM adapter per file path in a globalThis-backed map for idempotency, but never checked whether the cached adapter had since been closed (e.g. by gracefulShutdown/resetDbInstance racing a reload); reusing that dead handle made the very next query throw sql.js's own bare string "Database closed" (not an Error) straight out of instrumentation-node.ts's previously-unguarded ensureDbInitialized() call, and Next.js's internal registerInstrumentation() wrapper unconditionally does err.message = ... on whatever register() rejects with — assigning .message on a primitive string throws in strict mode, so the secondary TypeError is what actually crashed the process. Fixed in two parts: preInitSqlJs() now evicts a closed cached adapter and creates a fresh one instead of returning it; a new ensureDbReadyForBoot() wraps the DB-init call, normalizes any non-Error throw via normalizeBootError(), and retries once specifically for a transient "database closed" message (now succeeding against the fresh adapter) before re-throwing anything else as a real Error. Regression guard: tests/unit/instrumentation-database-closed-6560.test.ts.

  • fix(api): POST /api/keys no longer hangs 20–90+ seconds on a fresh install, even with valid auth (#6570) — cloudEnabled defaults to true in src/lib/db/settings.ts::getSettings() on any install with no persisted settings row (i.e. every fresh install), so the create-key handler's unconditional await syncKeysToCloudIfEnabled() always attempted a real outbound fetch() to CLOUD_URL via syncToCloud(); when that endpoint is unset/unreachable/slow, the HTTP response blocked until the request settled or timed out — unlike sibling routes (POST /api/keys/:id/regenerate, GET /api/combos), which never touch this side effect at all. src/app/api/keys/route.ts now dispatches syncKeysToCloudIfEnabled() fire-and-forget (void) instead of awaiting it; its internal try/catch already logs failures, so cloud sync still runs, it just no longer blocks the response. Regression guard: tests/unit/api-keys-create-no-hang-6570.test.ts (asserts the route resolves in well under 2s even when the Cloud-sync fetch() is stubbed to never settle) + updated tests/integration/api-keys.test.ts (the pre-existing "triggers cloud sync"/"still succeeds when cloud sync fails" tests, which previously hung indefinitely on this exact path, now await the fire-and-forget sync tick before asserting).

  • fix(api): editing any existing OpenAI Codex provider connection in the dashboard returned "Invalid request" and the edit could never be saved (#6562) — createProviderConnection() (src/lib/db/providers.ts) auto-increments a new connection's priority to MAX(priority)+1 per provider with no upper bound, and OAuth-imported connections (Codex codex-auth/import / import-bulk, up to 50 accounts per call, callable repeatedly — the standard Codex bulk-account-rotation workflow) never pass through createProviderSchema's Zod validation at all, so nothing ever capped that value; EditConnectionModal's handleSubmit always resends the connection's current priority unchanged on every save, and updateProviderConnectionSchema capped priority/globalPriority at max(100) — a UI-only ceiling the create path never enforced — so the first edit of any connection whose priority had already grown past 100 (routine once a Codex account count exceeds 100) failed validation regardless of which field the user actually changed. Raised the ceiling to max(100_000) on both fields — still bounded (a genuinely out-of-range value is still rejected), just wide enough to accept priorities the app itself already produces. Regression guard: tests/unit/codex-connection-edit-6562.test.ts (a Codex OAuth connection whose priority already exceeds the old 100 cap now validates + persists on edit; a control payload with a still-genuinely-invalid priority is still rejected with "Invalid request").

  • mimocode: rotate accounts on MiMoCode's rate-limit-style 400s (body-classified) instead of failing on the first account; malformed 400s still fail fast with the real upstream error (#6648 — thanks @pizzav-xyz)

  • fix(cli): compression CLI REST fallback now reads/writes the canonical defaultMode field (surfaced as strategy) instead of a nonexistent engine key, and table output renders nested objects as JSON instead of [object Object] (#6571 — thanks @charleszolot)

  • fix(providers): web-cookie providers without a providerRegistry.ts entry (lmarena, gemini-business, poe-web, venice-web, v0-vercel-web) now report unsupported: true instead of silently "OK" (#6309) — validateWebCookieProvider() (src/lib/providers/validation.ts) previously required a registry entry and returned "Provider not found in registry" for these; a fallback to WEB_COOKIE_PROVIDERS[provider].website was proposed, but live verification showed the ${website}/models probe does not reliably signal session validity for these providers (redirects/SPA 200s regardless of cookie validity — e.g. lmarena's real API is arena.ai, not lmarena.ai; Poe's real endpoint is a GraphQL POST, not a REST /models), so it would report an expired or garbage cookie as valid. Until each provider has a verified, side-effect-free auth probe against its real API host, the fallback now returns unsupported (no network call) instead of a false positive. Regression guard: tests/unit/web-cookie-validation-fallback.test.ts. (thanks @oyi77)

  • fix(api): POST /api/middleware/hooks and PUT /api/middleware/hooks/[name] no longer leak raw internal error messages in their 500 responses (#6645 — thanks @chirag127) — both catch blocks returned error?.message directly (Hard Rule #12), which could surface internal SQLite path fragments on a DB failure; both now route through sanitizeErrorMessage() from open-sse/utils/error.ts. Regression guard: tests/unit/middleware-hooks-error-sanitization.test.ts.

  • fix(docker): compile better-sqlite3 for the server Docker image (Dokploy/self-hosted builds) via a direct node-gyp rebuild inside node_modules/better-sqlite3, instead of npm rebuild better-sqlite3 (#6700) — the builder stage installs dependencies with npm ci --ignore-scripts (deliberate: closes the supply-chain surface where a transitive dep's install script runs arbitrary code) and re-enables the native build for the one package that needs it; npm rebuild <pkg> re-runs that indirectly through the package's own install script, which under npm 11 depends on npm's script-allowlist machinery correctly re-enabling it — some self-hosted build environments (e.g. Dokploy) hit a broken/mismatched native binding through that indirection. Invoking node-gyp rebuild directly bypasses npm's script-running layer entirely and is deterministic regardless of npm version. Regression guard: tests/unit/dockerfile-better-sqlite3-node-gyp-6700.test.ts. (thanks @nowhats-br)

  • fix(providers): the Cloudflare relay Worker deploy fix in #6416/#6618 still failed uploads in practice — it changed the multipart Content-Type but kept the emitted worker source as an ES module (export default { fetch(...) }) with main_module metadata; Cloudflare's Workers upload API parses a plain application/javascript script part as Service Worker syntax regardless of the main_module metadata field, and main_module requires the script to actually be an ES module (top-level export), so the mismatch still rejected the upload (#6496). buildCloudflareWorkerScript() (src/lib/proxyRelay/cloudflareWorkerScript.ts) now emits Service Worker syntax (addEventListener("fetch", ...), no top-level export) and the upload metadata uses body_part instead of main_module. Regression guard: tests/unit/relay-deploy-5128.test.ts (asserts the emitted script has no export default, registers a fetch listener, and the upload metadata carries body_part/omits main_module; also proves the inlined isPrivateHostname() SSRF guard still rejects bracketed IPv6 loopback/ULA hosts like [::1]/[fd00::1] after the script-body rewrite). (thanks @SeaXen)

  • fix(providers): ChatGPT Web (chatgpt-web) responses rendered raw ChatGPT UI citation markup — private-use marker tokens (e.g. citeturn0search0) and url… inline-link markers — instead of real Markdown links, since these only ever get resolved client-side by chatgpt.com's own JS using message.metadata.content_references (#6635) — cleanChatGptText() now resolves content_references (grouped webpages, footnote sources, inline webpage/url mentions) into [label](url) Markdown links for both the streaming and non-streaming response builders, and for the GPT-5.5 Pro stream_handoff polled-answer path, falling back to stripping any marker that has no resolvable source instead of leaking the raw private-use bytes. The citation parsing/rendering logic was extracted into a new pure sibling module (open-sse/executors/chatgpt-web/citations.ts) to keep the executor under the frozen file-size cap. Regression guard: tests/unit/chatgpt-web-citations.test.ts (non-streaming citation resolution, streaming marker buffering across split SSE chunks, and the Pro-handoff polled-answer path). (thanks @Thinkscape)

  • fix(dashboard): the live-dashboard WebSocket descriptor handshake (GET /api/v1/ws?handshake=1) and the lightweight GET /api/health/ping liveness probe both 401'd for unauthenticated callers, even though both are metadata-only reads intended to be public (#6335) — clientApiPolicy required a bearer/dashboard-session before the WS route handler could even return its own wsAuth/protocol descriptor, and /api/health/ping was never added to PUBLIC_READONLY_API_ROUTE_PREFIXES despite its own docstring documenting it as "No auth required." clientApiPolicy.evaluate() now allows an anonymous {kind:"anonymous", id:"ws-handshake"} subject for GET/HEAD/OPTIONS on /api/v1/ws?handshake=1 (the route handler still performs its own real wsAuth/dashboard/API-key decision before opening the socket), and /api/health/ping is now in PUBLIC_READONLY_API_ROUTE_PREFIXES. Regression guard: tests/unit/authz/client-api-policy.test.ts (WS handshake allowed, including relative request URLs), tests/unit/public-api-routes.test.ts, and tests/unit/authz/classify.test.ts (/api/health/ping classified PUBLIC). (thanks @JxnLexn)

  • fix(providers): wire the Devin cloud-agent provider into the generic provider-page validator and static model catalog, matching the existing jules cloud-agent pattern (#6142)

  • fix(providers): honor a provider-level proxy assigned to no-auth providers like MiMoCode Free (#6272)

  • fix(api): merge tool_call continuation deltas that carry only id (no index) so tool-call arguments are no longer split/lost in request/response logs (#6276)

  • fix(providers): modernize the lmarena provider for the Arena.ai rebrand — route chat through arena.ai create-evaluation with Chrome TLS impersonation, seed a static Direct-chat Text/Search + Image catalog, and keep the lmarena/lma wire id for back-compat (#6280) — thanks @backryun

  • fix(providers): web-provider model discovery updated — qwen-web uses the slash-terminated models endpoint (avoiding a blocked 307 redirect), and kimi-web matches the current request shape (POST with bearer + kimi-auth cookie replay) with its catalog refreshed to the current non-agent models (#6308 — thanks @janeza2).

  • fix(logs): the request-log detail modal no longer reopens by itself after being closed — a stale in-flight detail refresh resolved after close and re-triggered the modal open state (#6323 — thanks @xz-dev).

  • fix(providers): update SenseNova Token Plan support — register the token-plan model ids/constants and adjust the SenseNova registry so token-plan accounts route correctly (#6330 — thanks @xz-dev).

  • fix(providers): give v0-vercel-web its own alias so its credentials are detected (#6343)

  • fix(providers): route AgentRouter key validation through the CC wire image so a valid key no longer 403s as "Invalid API key" (#6377)

  • fix(db): stop legacy log-archive migration from deleting the live app-logger directory and crashing startup on a stat/stream race (#6401, #6799)

  • fix(docs): document Turbopack build memory tradeoff and OMNIROUTE_USE_TURBOPACK=0 webpack fallback for RAM-constrained machines (#6409)

  • fix(compression): surface silently-dropped stacked-pipeline steps (session-dedup, ccr) and stop the aggregate inflation guard from misfiring on a genuine no-op (#6479, #6480, #6491)

  • fix(providers): honor the max_token capability override in the reasoning-token-buffer output cap (#6524)

  • fix(dashboard): the onboarding tier-flow diagram rendered broken — its SVGs lived in the repo-root images/ (not a served path); moved to public/images/ so Next.js serves them (#6538 — thanks @ianriizky).

  • fix(routing): the auto combo's no-auth candidate pool now honors a disabled provider connection's own isActive=false (the toggle on the main Providers grid card), not just the separate global blockedProviders setting — disabling opencode/mimocode/etc. via the grid toggle no longer leaves it in rotation (#6557).

  • fix(sse): server-tool literal names (e.g. web_search) are preserved in message history and tool_choice instead of being namespaced/rewritten, so follow-up turns referencing those tools keep working (#6586 — thanks @MikeTuev).

  • fix(api): recognize OpenRouter reasoning/reasoning_details in non-streaming OpenAI-to-Claude conversion (#6623)

  • fix(db): share one in-flight sql.js load across concurrent preInitSqlJs() callers to stop the boot-time thundering-herd re-decode of the whole database file (#6628)

  • fix(db): unwrap lone named-parameter objects before sql.js stmt.bind() so @/:/$-style named placeholders bind correctly instead of throwing "Wrong API use" (#6802)

  • fix(db): break probe-failed/restore loop on large storage.sqlite (#6632 — thanks @KooshaPari).

  • fix(ci): exclude check-test-masking.test.ts's own tautology fixtures from the diff-based test-masking gate and recognize in validate-release-green's failure-line detector (#6634)

  • fix(routing): recognize Kimi-style "exceeded model token limit" 400 as context overflow so combo fallback continues to the next target (#6637)

  • fix(cli): Claude Code installed via WinGet is now detected on Windows (the WinGet install path was missing from the binary lookup) (#6647 — thanks @enjoyer-hub).

  • fix(providers): removed obsolete/defunct providers from the catalog (glhf, kluster, cablyai, inclusionai) (#6675 — thanks @backryun).

  • fix(sse): requests rejected before handleChatCore (circuit-breaker/cooldown gate or combo with all targets exhausted) are now recorded in usage_history too, so a key whose traffic was entirely gate-rejected no longer shows "zero requests" in the per-API-key usage counter (#6698).

  • fix(sse): unwrap bare {function:{…}} tools so OpenAI-shape clients no longer have tools silently dropped in Claude translation. (thanks @samir-abis)

  • fix(oauth): stop merging distinct Codex OAuth logins that share an email but lack a verifiable account id, preventing silent token overwrite. (thanks @lucasjustinudin)

  • fix(codex): detect "model at capacity"/overloaded errors embedded in a 200-OK SSE stream and surface them as a real error so account fallback rotates, instead of passing them through as a successful response. (thanks @ryanngit)

  • fix(volcengine): clamp max_tokens to the VolcEngine Ark endpoint cap for the Kimi model so oversized values no longer 400. (thanks @whale9820)

  • fix(antigravity): surface aborted/malformed Gemini tool calls (e.g. MALFORMED_FUNCTION_CALL) as an explicit non-end_turn finish reason instead of a silent clean completion. (thanks @anhdiepmmk)

  • fix(routing): the reasoning-token headroom buffer clamps to the model's explicit output cap instead of inflating past it, and getExplicitModelOutputCap falls through to the registry/spec cap when a synced capability row exists without a numeric limit_output (#6714) — thanks @xz-dev

  • fix(api): omniroute health (and health components/health watch) returned Error: HTTP 404 (#6677) — bin/cli/commands/health.mjs called apiFetch("/api/health", ...), a route that was moved to GET /api/monitoring/health (src/app/api/monitoring/health/route.ts) without updating the CLI; src/app/api/health/ on disk only has degradation/route.ts and ping/route.ts, no top-level handler. runHealthCommand()/runHealthComponentsCommand() now call /api/monitoring/health and read its actual payload shape (activeConnections, circuitBreakers: {open, halfOpen, closed}, memoryUsage) instead of the old, nonexistent requests/breakers/cache/memory fields. Regression guard: tests/unit/cli-health-monitoring-route.test.ts.

  • fix(startup): webpack build broke on case-insensitive filesystems (macOS APFS default, Windows) with a casing-collision warning plus "not exported" errors in StudioConfigPane.tsx/ChatTab.tsx (#6584) — src/app/(dashboard)/dashboard/playground/components/ReasoningControls.tsx (the component) and reasoningControls.ts (the utils module) shared the same lower-cased stem in the same directory, and two importers used the extensionless form from "./reasoningControls", the exact resolution path that becomes ambiguous once casing is folded. Renamed the utils module to reasoningControlUtils.ts (no collision) and updated the 3 import sites. Regression guard: tests/unit/case-collision-6584.test.ts (scans src//open-sse/ for any same-directory, case-only filename collision). (#6584)

  • fix(build): Turbopack production build emitted an "Overly broad patterns can lead to build performance issues" warning per entry point importing src/lib/agentSkills/generator.ts (603 warnings reported on v3.8.46, up from 379 on v3.8.45) (#6582) — generator.ts's outputBase is built as path.isAbsolute(outputDir) ? outputDir : path.join(process.cwd(), outputDir), where outputDir is a runtime function parameter, not a compile-time literal, so Turbopack's build-time file-tracing analyzer can't statically narrow the several dynamic readdirSync/rmSync/readFileSync/writeFileSync call sites a few lines below and falls back to a project-wide glob; #6366's commit message claimed to "anchor the base path with a literal" but the shipped code never did. Since this fs access is legitimate and bounded (skills/<id>/SKILL.md, ~48 known IDs), next.config.mjs's turbopack.ignoreIssue (Next.js 16.2+) now suppresses this specific, known-benign diagnostic, mirroring the existing webpack.ignoreWarnings/isNextIntlExtractorDynamicImportWarning precedent already in the same file for the webpack path. Regression guard: tests/unit/next-config.test.ts (asserts the turbopack.ignoreIssue rule shape targeting src/lib/agentSkills/**).

  • fix(providers): Codex Desktop requests to gpt-5.3-codex-spark failed with [400]: Tool 'image_generation' is not supported with gpt-5.3-codex-spark, even on paid-plan accounts (#6651) — CodexExecutor.transformRequest (open-sse/executors/codex.ts) only dropped the Codex Desktop-injected image_generation hosted tool when isCodexFreePlan() matched the account's plan, with no awareness that Spark-scope Codex models reject image_generation upstream regardless of plan. dropImageGeneration now also drops it when getCodexModelScope(model) === "spark" (the existing Spark classifier from open-sse/config/codexQuotaScopes.ts), independent of account plan. Regression guard: tests/unit/codex-spark-image-generation.test.ts (thanks @alltomatos for independently catching and fixing it via #6819).

  • fix(providers): the provider quota card's weekly/session bars re-sorted by remaining percentage instead of staying in a fixed, deterministic order (#6687) — QuotaCardExpanded.tsx's sortQuotasByRemaining() (added in #5977) was applied unconditionally via useMemo(() => sortQuotasByRemaining(quotas), [quotas]), undoing the deterministic CODEX_QUOTA_ORDER/GLM_QUOTA_ORDER window order quotaParsing.ts's sortCodexOrder()/sortGlmOrder() (added in #6336) already established for Codex and the GLM family — since #6336 never touched QuotaCardExpanded.tsx, the two orderings never composed, so e.g. a Codex session window with less headroom than weekly rendered after it instead of staying first. A new hasFixedQuotaOrder() (quotaParsing.ts) and resolveQuotaDisplayOrder() (QuotaCardExpanded.tsx) now skip the remaining-% re-sort for providers with a fixed window order, threading providerId from QuotaCard.tsx through to the display layer; every other provider still gets the remaining-% sort. Regression guard: tests/unit/quota-card-expanded-fixed-order-6687.test.ts.

  • fix(i18n): pt-BR was missing 194 UI keys present in en.json — a real, silent data-sync gap, not covered by any duplicate/mislabeled #6694 (that issue's 9 providers.* keys are disjoint, present-but-untranslated sentinels caused by a separate providerText() fallback bug) (#6695) — scripts/i18n/sync-ui-keys.mjs (which mirrors newly-added en.json keys into every locale) wasn't re-run after recent en.json additions, and the CI i18n:check-ui-coverage gate only fails a locale below an 80% threshold, so pt-BR stayed green at 93.8% coverage despite the gap. Backfilled all 194 missing keys into src/i18n/messages/pt-BR.json (translated to Brazilian Portuguese, no leftover __MISSING__ markers) via npm run i18n:sync-ui -- --locale=pt-BR + manual translation. Regression guard: tests/unit/i18n-pt-br.test.ts (new case asserting full en.jsonpt-BR.json key parity, so a future drift fails a fast unit test instead of silently degrading the coverage percentage).

  • fix(startup): omniroute --mcp crashed at Node ESM link time with ERR_MODULE_NOT_FOUND for ioredis on installs where the published MCP bundle didn't happen to have ioredis rescued from a parent node_modules (#6559) — src/shared/utils/rateLimiter.ts had a top-level static import Redis from "ioredis"; that module is only ever reached via a lazy await import(...) several call-sites deep in the MCP tool chain, but esbuild's --packages=external bundling of the MCP server (scripts/build/prepublish.ts Step 8.5) still hoisted rateLimiter.ts's own static import into a real top-level ESM import in the compiled dist/open-sse/mcp-server/server.js, forcing Node to resolve ioredis at module-link time — before any --mcp startup code runs — and ioredis is not guaranteed to ship in the MCP-only bundle's node_modules. getRedisClient() now lazily imports ioredis on first use (matching the established soft-dependency pattern in src/lib/quota/redisQuotaStore.ts) while still throwing synchronously when Redis isn't configured. Regression guard: tests/unit/build/mcp-bundle-no-eager-ioredis.test.ts (bundles the real MCP server entrypoint with the exact publish-time esbuild flags and asserts no top-level static ioredis import remains, while the pre-existing lazy await import("ioredis") in redisQuotaStore.ts stays intact).

  • fix(providers): Kiro sent the adaptive-thinking additionalModelRequestFields envelope for claude-sonnet-4.5/claude-haiku-4.5, which Kiro/CodeWhisperer rejects upstream with a raw [400]: additionalModelRequestFields is not supported for this model (#6576) — buildKiroPayload() (open-sse/translator/request/openai-to-kiro.ts) gated the field on the generic Anthropic-API supportsReasoning() capability flag, which is true for both models on Anthropic's direct API but does not reflect what Kiro's CodeWhisperer backend actually accepts; only claude-sonnet-5 is confirmed adaptive-thinking-capable there. A new Kiro-specific allowlist (supportsKiroAdaptiveThinking() in open-sse/translator/request/openai-to-kiro/adaptiveThinking.ts) now gates the envelope instead. Regression guard: tests/unit/repro-6576-kiro-thinking-unsupported-model.test.ts.

  • fix(translator): Cursor's local Subagent tool call is no longer rejected with cloud_base_branch may only be specified when environment equals cloud — the Responses→Chat tool-arg cleanup (stripEmptyOptionalToolArgs) was scoped to Claude Code's Read tool only, so Cursor's Subagent tool passed through with the cloud-only cloud_base_branch: "" (Cursor treats an empty string as "specified" and rejects the call before starting the local subagent). The cleanup now covers an allowlist of Read + Subagent; arbitrary tools are still left untouched (empty strings/arrays can be valid payloads for them). Regression guard: tests/unit/openai-responses-subagent-strip-2446.test.ts. (thanks @like3213934360-lab)

  • fix(translator): GLM 5.2 (and other OpenAI-compatible upstreams that stream a tool call's id and function.name in separate SSE chunks) no longer produce an empty tool name / No such tool available: error through the Claude /messages path — the openai-to-claude streaming translator emitted content_block_start immediately on the id-only chunk with an empty name, and the Claude SSE protocol cannot patch a block after it is emitted, so the later name-only chunk was silently dropped. It now defers content_block_start until the tool name arrives (falling back to starting the block when arguments arrive first), so the emitted tool_use always carries the real name. Regression guard: tests/unit/openai-to-claude-glm-split-tool-name-2077.test.ts. (thanks @itiwant)

  • fix(resilience): OmniRoute didn't respect an exhausted Ollama Cloud (or any other apikey-category provider) quota — it retried the account seconds later instead of waiting out the real reset window (#6638) — shouldPreserveQuotaSignalsFor429()/checkFallbackError() (open-sse/services/accountFallback.ts) only applied body-text quota classification (daily/monthly/weekly quota-exhausted detection) to OAuth-category providers; apikey-category 429s (Ollama Cloud, OpenAI, etc.) always fell through to the generic short rate-limit cooldown regardless of what the error body said, and parseRetryFromErrorText() also had no support for day-granularity reset hints ("Your quota will reset in 3 days.") — only Xh/Ym/Zs combos. An explicit quota-exhausted signal in the body (looksLikeQuotaExhausted()) now overrides the apikey-category default via the new shouldPreserveQuotaSignals() (open-sse/services/quotaResetParsing.ts), and parseDayGranularityResetMs() parses whole-day reset countdowns so the real multi-day window is honored instead of a few seconds of backoff. Regression guard: tests/unit/issue-6638-ollama-quota.test.ts + 2 aligned tests/unit/account-fallback-service.test.ts cases that previously asserted the buggy rate_limit_exceeded/undefined-dailyQuotaExhausted behavior for apikey-provider quota text.

  • fix(resilience): a combo step "pinned" to one fingerprint account (mimocode/mcode/opencode multi-account providers) never actually resolved to that account, so it couldn't fail over when the pinned account was depleted (#6696, relates #6612) — the combo builder UI encodes an account pin as a composite connectionId (${rowId}|fp|${fingerprint}, src/lib/combos/builderOptions.ts), but expandTargetsByFingerprints() (open-sse/services/combo/fingerprintExpansion.ts) looked that composite string up directly in connectionById (keyed by real DB row ids), got undefined, and passed the target through unchanged, still carrying the bogus composite id — so downstream credential resolution could never match it either. expandTargetsByFingerprints() now splits the |fp| composite id back into the real connection row id + the pinned fingerprint (new splitFingerprintPin() helper) before any lookup, resolving the target to the real connectionId (with the pinned fingerprint carried on the new pinnedFingerprint field) instead of the inert composite string. Regression guard: tests/unit/combo-fingerprint-pin-6696.test.ts.

  • fix(api): Responses passthrough emitted event-only SSE frames (no data: line) for every dropped commentary event, breaking the OpenAI Python SDK's sse.json() parser (#6561), follow-up to #6199/#6232 — the commentary-drop continue; branches in open-sse/utils/stream.ts skipped the data: line for a dropped commentary event but never cleared the already-buffered event: line for that same frame, so the next blank line flushed the stale event: line alone. Both drop sites now call clearPendingPassthroughEvent() before continue, discarding the buffered prefix along with the dropped payload; the commentary-drop decision itself was extracted into a new open-sse/utils/responsesCommentaryDrop.ts so the fix does not grow the frozen stream.ts. Regression guard: tests/unit/responses-commentary-event-frame-6561.test.ts (realistic event:\ndata:\n\n frames — the existing #6199 test only used bare data: lines and never exercised this path).

View originalPermalink
How v3.8.48 went

v3.8.47

Added 16
  • Langfuse observability plugin
  • Context requirements config for per-target filtering in combos
  • Icons for 46 providers that were missing images
  • Shorthand proxy formats and protocol header mode for bulk import
  • OpenVecta AI inference gateway provider
  • Traditional Chinese (zh-TW) localization for frontend and CLI
Changed 2
  • Vendored GCF (Headroom) codec updated to spec v3.2 with nested flattening
  • Codex bulk-import endpoint now accepts 9router's camelCase account export format alongside snake_case
  • 9router Codex import: the Codex bulk-import endpoint (POST /api/oauth/codex/import) now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt + nested providerSpecificData), not just snake_case — normalizeCodexImportRecord maps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged (#6665) — thanks @deadcoder0904. Regression guard: tests/unit/codexBulkImport.test.ts (9router camelCase record, pre-supplied providerSpecificData without an id_token, snake_case-not-overridden, and a full {accounts:[...]} flatten).
✨ New Features
  • feat(plugins): Langfuse observability plugin. (#6577 — thanks @chirag127)

  • feat(combo): context requirements config for per-target filtering in combos. (#6907 — thanks @oyi77)

  • feat(providers): icons for 46 providers that were missing images. (#6926 — thanks @oyi77)

  • feat(compression): vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). (#6838 — thanks @blackwell-systems)

  • feat(proxy): shorthand proxy formats + protocol header mode for bulk import. (#6867 — thanks @growab)

  • feat(provider): OpenVecta AI inference gateway. (#6833 — thanks @hajilok)

  • feat(i18n): Traditional Chinese (zh-TW) localization for frontend and CLI. (#6320 — thanks @lunkerchen)

  • feat(xai): route xAI clients to Grok's native /v1/responses endpoint. (#6709 — thanks @diegosouzapw)

  • feat(routing): per-model web-search/web-fetch interception rules. (#3384, #6814 — thanks @diegosouzapw)

  • feat(release): changelog.d/ fragments — eliminates the CHANGELOG merge-storm cascade. (#6783 — thanks @diegosouzapw)

  • feat(quality): validate-release-green --full-ci reproduces the entire ci.yml static gate set locally. (#6583 — thanks @diegosouzapw)

  • feat(dashboard): sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (src/shared/components/Sidebar.tsx) filters nav sections/groups/items client-side by label as you type, reusing the existing common.search/common.noResults i18n keys (zero new locale edits) and the shared Input icon="search" pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted into filterSidebarSectionsByQuery() (src/shared/utils/sidebarSearch.ts) for isolated unit testing. Regression guard: tests/unit/sidebar-search-filter.test.ts, src/shared/components/Sidebar.search.test.tsx. (#4013 — thanks @crochabe-cyber)

  • feat(combo): auto/* combos gain a strict budget-cap fallback policy — X-OmniRoute-Budget-Fallback: strict (or the persisted config.budgetFallback: "strict") makes an over-budget request fail fast with HTTP 402 instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (cheapest) preserves existing behavior. Builds on the existing X-OmniRoute-Budget/X-OmniRoute-Mode per-request controls (#6023/#6024/#6025), consolidated into resolveRequestAutoControls(). Regression guard: tests/unit/auto-combo-budget-fallback-3470.test.ts. (#3470)

  • Provider/model param filters: config-driven parameter denylist/allowlist per provider/model with auto-learn from upstream 400s (#6649 — thanks @ThongAccount, closes #6625)

  • Per-combo reasoning token buffer toggle: the combo builder now exposes an explicit checkbox for the #3587 reasoning-model max_tokens buffer, defaulting to the existing enabled behavior, so a combo can opt out without hand-editing raw JSON config (#6702 — thanks @xz-dev)

  • feat(dashboard): 9router-parity Routing Strategy settings card on Settings → Routing, plus a per-provider account-routing override on the provider detail page (#6678) — surfaces the existing account round-robin / sticky-limit knobs and adds a new combo-level sticky round-robin (comboStickyRoundRobinLimit, resolved via resolveComboStickyRoundRobinLimit() — per-combo → global combo sticky → account sticky cascade) so combo targets can batch calls per target the same way account fallback already does. A new providerStrategies setting (Zod-validated map, src/shared/validation/settingsSchemas.ts) lets a specific provider override the global fallbackStrategy/stickyRoundRobinLimit without touching the account-wide default, wired into getProviderCredentials() (src/sse/services/auth.ts) ahead of the global fallback. Regression guard: tests/unit/combo-rr-sticky-9router.test.ts, tests/unit/settings-ui-layout-static.test.ts. (thanks @SeaXen)

  • feat(icons): provider logos now resolve local SVG assets first for faster rendering, with a 5-tier fallback chain — local SVG → @lobehub/icons React components → thesvg.org CDN (external SVG for unknown providers) → local PNG → generic AI icon — replacing the previous LobeHub-first order. Adds dozens of first-party provider SVGs and migrates several bitmap logos (continue/copilot/cursor/deepgram/heroku/openclaw/ovhcloud) from PNG to SVG. Regression guard: tests/unit/ui/ProviderIcon-icon-url.test.tsx. (#6317 — thanks @hamsa0x7)

  • Skill Collector CLI detection: new GET /api/skills/collect/detect + POST /api/skills/collect/install (and the cli-skill-collector agent skill) detect which coding CLIs (Claude Code, Codex, Cursor, Copilot, Cline, Hermes, OpenCode, etc.) are installed locally via getCliRuntimeStatus(), match them against GitHub agent-skill repos, and plan an install path per tool — replacing the standalone Skill Collector Python app. Both new routes and GET/POST /api/github-skills now require management auth (requireManagementAuth()) and are loopback-gated (LOCAL_ONLY_API_PREFIXES + SPAWN_CAPABLE_PREFIXES) since the detect route spawns a child process per candidate CLI tool (Hard Rules #15 + #17). The omniroute_github_skills_install MCP tool now reports the honest action: "planned" instead of "installed", matching the REST route (#6294 — thanks @Moseyuh333)

  • ClinePass dual-auth: ClinePass now offers both sign-in methods on its dashboard page — OAuth (reusing the Cline WorkOS flow) as the primary "Connect" path, or a pasted BYOK API key via "Manual API key", instead of only the API-key-only provider shipped in #5942. The registry alias was aligned to cp (matching the OAUTH_PROVIDERS catalog alias) so <alias>/<modelId> routing resolves correctly, the OAuth refresh dispatch now routes clinepass to the shared Cline refresh flow, and the duplicate API-key-only catalog entry was removed to keep ClinePass listed once. Regression guard: tests/unit/clinepass-provider.test.ts. (#6126 — thanks @hajilok)

  • feat(oauth): Kiro/Amazon Q auto-import now supports enterprise External IdP ("Your organization") logins via Microsoft Entra/Okta/Auth0/OneLogin/Ping/Google/Cognito — these org-issued tokens are not AWS SSO tokens (no aorAAAAAG-prefixed refresh token) and can't refresh through the AWS OIDC/Kiro-social path, so tryAwsSsoCache() now detects them (authMethod/provider === "externalidp") and refreshes via the org IdP's own tokenEndpoint (public-client OAuth2 refresh grant, no client secret), persisting TokenType: EXTERNAL_IDP gating so the runtime executor sends the header the AWS CodeWhisperer API requires for these accounts; tokenEndpoint is SSRF-guarded against an HTTPS + known-IdP-host-suffix allowlist. (#6363 — thanks @artickc)

  • Kiro long-lived API key auth: new /api/oauth/kiro/api-key route + KiroService.validateApiKey let a Kiro account be linked with a long-lived AWS CodeWhisperer/Kiro API key instead of the interactive OAuth device flow, with live per-account model discovery (ListAvailableModels, 5-minute cache) layered over the existing static registry fallback (#6587 — thanks @strangersp)

  • Chaos Mode: multi-model parallel/collaborative task execution — dispatches a task to every active provider connection at once (parallel) or chains outputs sequentially so each model builds on the previous one's answer (collaborative), configurable via Dashboard → Chaos Mode (GET/PUT/DELETE /api/chaos/config) and gated per-API-key via a new chaosModeEnabled permission (opt-in — disabled by default globally and per key). POST /api/chaos/run (dashboard session) and POST /api/skills/collect/chaos (external Bearer-token) delegate to a shared executeChaosRun() engine (src/lib/chaos/chaosExecutor.ts) that dispatches in-process via the established synthetic-Request/route-handler pattern (no network hop, no hardcoded port), with a concurrency cap (max 10 parallel), configurable max_tokens (256–128k), a clear error when stream is requested, and collaborative-chain info (provider order + input size). Fixes external Bearer-auth bypass and stale config-cache leakage. Regression guard: tests/unit/chaos-config.test.ts, tests/unit/chaos-executor.test.ts, tests/unit/chaos-api-routes.test.ts. (#6728 — thanks @Moseyuh333)

  • feat(cli): 2 new CLI tool integrations on Dashboard → CLI Tools — omp (Oh My Pi) and letta — each with binary detection, config apply/reset, and a settings card following the existing tool-card pattern. Both settings routes shell out to which omp/which letta to detect the local install, so they're loopback-gated (LOCAL_ONLY_API_PREFIXES, Hard Rules #15/#17) in addition to the shared requireCliToolsAuth() management-auth guard every cli-tools route requires, and route errors through sanitizeErrorMessage(); src/lib/db/omp.ts isolates the omp CLI's own local SQLite reads behind parameterized queries. (Note: the original PR also proposed pi, codewhale, and jcode integrations — those three had already shipped via a separate PR by the time this one was reconciled, so only omp+letta landed here.) Regression guard: tests/unit/db/omp.test.ts, tests/unit/cli-tools-auth-hardening.test.ts, tests/integration/cli-settings-omp.test.ts, tests/integration/cli-settings-letta.test.ts. (#6318 — thanks @hamsa0x7)

  • feat(providers): custom models now support a manual Context Window Override so an operator can correct a provider's misreported context length (e.g. reports 1M when the real limit is 128K) instead of the model silently getting dropped from combo routing once the wrong value lands in the catalog (#4125 — thanks @rucciva). Reuses the existing Feature-5004 model_context_overrides table (source: "manual") — already the priority-0 source getModelContextLimit() (the function combo's context-window filter calls) reads ahead of the models.dev/registry/static catalog — so no new resolver logic was needed, only the missing write path: PUT /api/provider-models now accepts an optional contextWindowOverride (number to set, null to clear), GET surfaces the current value back on each custom-model row, and the provider detail page's custom-model edit form gained a Context Window Override field + badge. Regression guard: tests/unit/provider-models-context-window-override-4125.test.ts.

  • fix(providers): register OpenRouter as a rerank provider so openrouter/cohere/rerank-* models resolve instead of erroring Invalid rerank model (#6574 — thanks @rafpigna)

  • fix(api): HEAD requests no longer hang until client timeout on any route — valid, unknown, authed, or unauthed (#6400), broader follow-up to the route-specific #6517 (/v1/models). Root cause: Next.js 16's App Router route-handler pipeline (next/dist/server/send-response.js) correctly skips piping a Response body for HEAD, but its page-rendering pipeline (next/dist/server/pipe-readable.jspipeToNodeResponse, used for every app-router page/layout render — including the not-found boundary any unmatched path falls through to) has no such check and always streams the full rendered body regardless of method; combined with Node's default keep-alive framing this left some clients unsure whether the (implicitly bodyless) HEAD response had actually finished. A new scripts/dev/head-response-guard.cjs, wired into both the dev/start custom server (scripts/dev/run-next.mjs) and the packaged standalone server (scripts/dev/standalone-server-ws.mjs) at the same tier as the existing http-method-guard.cjs/peer-stamp.mjs wrappers, discards any body bytes written for a HEAD request and forces Connection: close once .end() is called — independent of route existence or auth state, satisfying RFC 9110 §9.3.2. Regression guard: tests/unit/head-request-closes-6400.test.ts.

  • feat(dashboard): Provider Quota page (Dashboard → Quota) fills horizontal whitespace before stacking vertically (#3520) — QuotaCardGrid previously stacked every provider group in a single vertical flex flex-col, and each group's own card grid didn't go multi-column until md (grid-cols-1 md:grid-cols-2 xl:grid-cols-3 2xl:grid-cols-4). Provider groups now flow into a 2-column CSS multi-column layout on very wide (2xl) screens instead of an unconditional vertical stack, and each group's card grid starts at 2 columns immediately (grid-cols-2 md:grid-cols-3 xl:grid-cols-4), reaching higher density sooner on narrower-but-not-mobile viewports. Regression guard: tests/unit/quota-card-grid-horizontal-layout.test.ts (thanks @gdevenyi).

  • feat(ws): the live-dashboard WebSocket server now auto-starts in-process (via instrumentation-node.ts) across every deployment mode — dev, production, Docker, Electron — with no separate sidecar script; the default WS port moved from 20129 to 20132 to avoid colliding with API_PORT in split-port setups, the deprecated OMNIROUTE_DISABLE_LIVE_WS env was consolidated into OMNIROUTE_ENABLE_LIVE_WS (default enabled), and the WS path is now derived from NEXT_PUBLIC_LIVE_WS_PUBLIC_URL's pathname (/live-ws fallback) (#6072 — thanks @ianriizky).

  • feat(compression): new omniglyph engine (context-as-image) — renders system prompt, tool docs, and dense history as compact PNG pages the model reads instead of text (~10× fewer tokens on the converted block; 59–70% end-to-end measured). Works stacked with RTK/Caveman (stackPriority: 90) or standalone (mode: omniglyph); restricted to Claude Fable 5 over the direct Anthropic route, fail-closed gates with skip:<reason> techniques, preview (stable: false, off by default) (#6556). Dependency bumped to omniglyph@^1.0.2 for upstream ReDoS fixes (#6661).

  • feat(sandbox): the skill sandbox gained a container-provider abstraction that auto-detects and uses the best native runtime per host — Apple Container (macOS 26+), WSL container (wslc.exe), OrbStack, Podman — instead of hardcoding docker run, removing the Docker Desktop requirement on macOS/Windows (#6611 — thanks @KooshaPari).

  • fix(sse): skip thinkingConfig for Gemma models on the OpenAI→Gemini path so OpenAI-shape clients no longer get a 400 from Vertex. (thanks @chy1211)

  • feat(xai): route xAI clients to Grok's native /v1/responses endpoint instead of the chat-completions bridge. (thanks @ryanngit)

  • feat(models): add a Settings → AI "Model Overrides" UI plus /api/model-capability-overrides CRUD and a model_capability_overrides table, letting operators set a manual max-output-token override per provider/model (#6727 — thanks @xz-dev).

  • feat(resilience): operator-configurable account rotation policy — a new rotationConfig layer lets operators tune how connections rotate on failure, wired into accountFallback (#6763 — thanks @artickc).

  • chore(cursor): add Grok 4.5 effort/fast model IDs (#6774 — thanks @andrewmunsell).

  • feat(codex): Codex provider model discovery now fetches the live catalog from chatgpt.com/backend-api/codex/models using Codex-shaped headers, falling back to a GitHub-hosted model manifest and then to the local static catalog when the live/GitHub sources are unavailable or return an unexpected shape — new src/app/api/providers/[id]/models/discovery/codex.ts (normalization, version-gating, merge/enrich against the local catalog) covered by tests/unit/provider-models-discovery-split.test.ts and tests/unit/provider-models-route-codex.test.ts (#6776 — thanks @JxnLexn).

  • feat(cursor): register the Opus 4.8, Fable 5, and Sonnet 5 model families for the Cursor Agent provider so the latest Claude/Fable model ids route correctly (#6779 — thanks @andrewmunsell).

  • Changelog fragments (changelog.d/): PRs now add their changelog entry as a new fragment file (changelog.d/{features|fixes|maintenance}/<PR>-<slug>.md) instead of editing CHANGELOG.md — two PRs never touch the same file, structurally eliminating the CHANGELOG-eat merge conflicts that forced a re-sync push + full CI re-run after every sibling merge (O(N²) CI runs in a merge-storm). scripts/release/aggregate-changelog.mjs (npm run changelog:aggregate) folds fragments into the living section at release reconciliation, and check:changelog-integrity now also validates fragment well-formedness. Regression guard: tests/unit/changelog-fragments.test.ts.

  • feat(proxy): add a latency-optimized proxy rotation strategy that ranks pool entries by measured round-trip latency, extending the existing round-robin/random/sticky proxy-pool selection (#6798 — thanks @iamraydoan).

  • feat(fusion): the fusion judge may now draw on its own knowledge and override the panel when every panel answer is wrong or incomplete, instead of being restricted to synthesizing only from panel output (#6804 — thanks @chirag127).

  • feat(dashboard): search box on the Playground's raw model <select> (#4086) — the shared ModelSelectModal (combo builder + CLI-code cards) already had search, but Playground's StudioConfigPane model dropdown stayed a flat unsearchable list, unusable once a provider like OpenRouter contributed 50+ models. Typing now filters the dropdown (Turkish-safe accent/case-insensitive match via matchesSearch), while the currently selected model always stays pinned in the list even if it no longer matches the query, so typing never silently swaps the active selection. Reuses the existing common.search i18n key (already translated in all 42 locales) — no new translation key needed. Regression guard: tests/unit/playground-model-selection-3731.test.ts (filterModelsByQuery), tests/unit/ui/playground-model-search-4086.test.tsx.

  • feat(usage): Antigravity/agy quota widget now surfaces the weekly window alongside the existing per-model 5-hour window (#4017) — the weekly limit isn't part of the per-model retrieveUserQuota response the fetcher already calls; it only appears in a separate, undocumented retrieveUserQuotaSummary RPC that groups models into families ("Gemini Models", "Claude and GPT models") with one weekly bucket per family. A new usage/antigravityWeeklyQuota.ts leaf fetches that RPC (cached, best-effort — a failure or unavailable RPC never breaks the existing per-model quotas) and parses the weekly bucket per group into gemini_weekly/claude_gpt_weekly quota entries, merged into the same quotas map the widget already renders generically. Regression guard: tests/unit/antigravity-weekly-quota-4017.test.ts (bucket parsing, the alternate quotaSummary-nested envelope, and end-to-end merge via getUsageForProvider).

  • feat(codex): Codex CLI compatibility shim — the Responses API response.created/response.in_progress/response.completed payloads now carry a model field (previously absent), and for Codex-CLI-originated requests it echoes the client-requested, effort-suffixed model id (e.g. gpt-5.5-xhigh) instead of the bare upstream id (gpt-5.5), so the Codex CLI status line/model button shows the active reasoning effort (#3697). openaiToOpenAIResponsesResponse (open-sse/translator/response/openai-responses.ts) now threads the upstream model into the Responses event objects; a new isCodexOriginatedHeaders() (open-sse/config/codexIdentity.ts, reusing PR #3481's originator/User-Agent detection) makes chatCore's existing opt-in echoRequestedModelName (#1311) model-echo pipeline fire automatically for Codex clients regardless of the setting, detected by request headers so it still applies when codex/gpt-5.5-xhigh is routed through a combo to a non-codex upstream; echoModelInObject/echoModelInSseLine (open-sse/services/responseModelEcho.ts) now also rewrite the nested response.model field the Responses API uses. /v1/models still returns models: [] for Codex (unchanged). Regression guard: tests/unit/codex-effort-model-echo-3697.test.ts.

  • Z.ai Web (free web-session provider): new zai-web web-cookie provider drives the free chat.z.ai consumer chat UI via a pasted browser session cookie, distinct from the existing API-key zai/glm/glm-cn/glmt providers (api.z.ai) — modeled on the doubao-web/venice-web cookie executors and the pre-existing chatglm-web credential requirement/token-extraction entries. ZaiWebExecutor (open-sse/executors/zai-web.ts) posts to chat.z.ai/api/chat/completions with the cookie forwarded both as Cookie and as Authorization: Bearer <token>, and normalizes both z.ai's internal delta_content/phase SSE envelope and a pass-through OpenAI-shaped choices[].delta frame into standard chat-completion chunks. Registered in WEB_COOKIE_PROVIDERS, WEB_SESSION_CREDENTIAL_REQUIREMENTS, the provider registry (zai-web entry, GLM-4.6/4.5/4.5V models), and tokenExtractionConfig.ts for in-app cookie capture. Regression guard: tests/unit/executor-zai-web.test.ts (16 tests — token extraction, frame parsing for both SSE shapes, streaming and non-streaming aggregation, error paths). (#4056)

  • feat(compression): update the vendored GCF codec behind the Headroom engine to spec v3.2 (nested flattening) (#6837). Homogeneous arrays whose rows carry nested objects/arrays now tabularize via >-prefixed path fields instead of a low-yield per-row fallback, so nested MCP tool-result rows (meta:{...}, tags:[...]) compact like flat rows. On representative shapes the update takes deeply-nested payloads the old codec left near-uncompressed from ~3% to ~32% vs JSON (k8s pods, cl100k_base), with shallow-nested rows seeing a small bump and flat arrays unchanged. Re-vendored from current gcf-typescript (zero runtime deps, MIT, SPDX-marked, generic-profile only); also folds in the [N]: inline-array quoting fix and canonical decimal formatting. Round-trip stays lossless (order-insensitive), and the decoder is hardened against prototype pollution (a __proto__/constructor path segment never mutates Object.prototype, and keys shadowing built-ins like toString now round-trip correctly instead of misparsing). Regression guard: tests/unit/compression/headroom-smartcrusher.test.ts (deep-nested + prototype-pollution cases).

  • feat(providers): Add GPT-5.6 support across OpenAI API, Codex, and ChatGPT Web, including Codex Max/Ultra efforts, VS Code metadata, Fast-tier credit accounting, curated live discovery, the Codex 0.144.1 client identity, and correct chat routing for models that also support image generation (#6862) - thanks @backryun

  • chore(providers): Align emitted Claude Code identity headers, bridge fingerprints, provider profiles, and documented defaults with claude-cli 2.1.207 (#6862) - thanks @backryun

🐛 Bug Fixes
  • fix(dashboard): the Proxy Registry settings page crashed at runtime (ReferenceError: poolLoaded/bulkImportOpen is not defined) — the #6625/#6909 hook-extraction refactors deleted 10 state declarations (poolLoaded, poolSaving, and the 8-member bulk-import family) while ~30 usages remained; all restored (caught by the release E2E; typecheck:core does not cover dashboard TSX — follow-up #7021). (thanks @diegosouzapw)

  • fix(combo): comboStickyRoundRobinLimit now defaults to inherit (null) instead of 1 — the literal default silently shadowed the documented batched round-robin rotation (stickyRoundRobinLimit: 3), flipping every round-robin combo to per-request alternation (#6678 follow-up, caught by the release CI). (thanks @diegosouzapw)

  • fix(ws): the standalone LiveWS startup script exited 0 without ever listening — its bootstrapped child re-spawned with the import-suppressor OMNIROUTE_ENABLE_LIVE_WS=0 and then honored it as an operator disable (#6072 follow-up, caught by the release CI). (thanks @diegosouzapw)

  • fix(api): compression PUT schema accepts every catalog engine. (#6792 — thanks @Pitchfork-and-Torch)

  • fix(compression): surface fallback reasons in the preview response. (#6461, #6519 — thanks @chirag127)

  • fix(providers): fail fast on an empty auto-combo pool instead of a 15s timeout. (#6458, #6546 — thanks @chirag127)

  • fix(compression): honor UI-toggled engines in the stackedPipeline dispatch + surface substitutions. (#6463, #6534 — thanks @chirag127)

  • fix(api): return 400 for missing/invalid messages before model resolution. (#6402, #6515 — thanks @chirag127)

  • fix(providers): enrich the model_cooldown 429 body with a retry_after ISO timestamp + credential count. (#6460, #6523 — thanks @chirag127)

  • fix(sse): default reasoning summary for effort-only Responses requests. (#6807 — thanks @rushsinging)

  • fix(sse): compression no-op treated as zero-savings, not inflation/silent-drop. (#6883 — thanks @chirag127)

  • fix(oauth): Trae OAuth client_id embedded via resolvePublicCred() (Hard Rule #11). (#6870 — thanks @chirag127)

  • fix(sse): combo path no longer trips the whole-provider breaker on a plain 429. (#6868 — thanks @chirag127)

  • fix(api): malformed JSON bodies now return 400 instead of 500. (#6871 — thanks @chirag127)

  • fix(fusion): judge selected from a surviving panel member when no explicit judge is configured. (#6869 — thanks @chirag127)

  • fix(sse): combo model lockout honors the parsed upstream quota reset. (#6863, #6866 — thanks @AgentKiller45)

  • fix(dashboard): logs detail modal no longer reopens on first close. (#6830 — thanks @MikeTuev)

  • fix(usage): honor xAI provider-reported exact cost. (#6711 — thanks @diegosouzapw)

  • fix(kiro): probe IdC region during profileArn discovery, cross-region (recovers #6099). (#6840 — thanks @diegosouzapw)

  • fix(antigravity): sanitize Cloud Code safety settings. (#6839 — thanks @diegosouzapw)

  • fix(translator): defer content_block_start until GLM streams the tool name. (#6730 — thanks @diegosouzapw)

  • fix(translator): strip empty cloud_base_branch from Cursor Subagent tool calls. (#6729 — thanks @diegosouzapw)

  • fix(antigravity): surface aborted Gemini tool calls off end_turn. (#6713 — thanks @diegosouzapw)

  • fix(volcengine): clamp Kimi max_tokens to the Ark endpoint cap. (#6712 — thanks @diegosouzapw)

  • fix(codex): surface capacity errors embedded in 200-OK SSE streams. (#6710 — thanks @diegosouzapw)

  • fix(sse): skip thinkingConfig for gemma models in openai→gemini translation. (#6708 — thanks @diegosouzapw)

  • fix(oauth): avoid bare-email dedup of Codex OAuth logins. (#6706 — thanks @diegosouzapw)

  • fix(sse): unwrap bare {function:{…}} tools in openai→claude translation. (#6704 — thanks @diegosouzapw)

  • fix(db): eliminate a redundant getApiKeyMetadata call in the embeddings route. (#6929 — thanks @oyi77)

  • fix(db): authType filter support in getProviderConnections. (#6946 — thanks @oyi77)

  • fix(i18n): the provider-detail (/dashboard/providers/[id]) visibility + free/paid model filter labels (showVisibleOnly, showHiddenOnly, freeFilterAll, freeFilterFreeOnly, freeFilterPaidOnly, hideAllModels, plus the currently-unused filterVisible/filterHidden/filterByVisibility) rendered as the literal __MISSING__:<english> sentinel in 15 locales, including pt-BR (#6694) — providerText() (providerPageHelpers.ts) checks t.has(key) before falling back to clean English, and t.has() returns true even when the stored value is the __MISSING__: sentinel scripts/i18n/sync-ui-keys.mjs writes when mirroring keys across locales, so the sentinel rendered verbatim instead of the fallback. Disjoint key set from #6290 (filterAll/filterActive/filterError/filterBanned/filterCreditsExhausted). All 9 keys now carry real translations across the 15 affected locale files (it, ja, ko, mr, ms, nl, no, phi, pl, pt, pt-BR, ro, ru, sk, sv). Regression guard: tests/unit/i18n-provider-visibility-filter-keys-6694.test.ts.

  • fix(cli): the dashboard's Claude Code CLI card could report "Not detected"/"Not installed" even when Claude Code was genuinely installed and previously used (#6701) — getCliRuntimeStatus() (src/shared/services/cliRuntime.ts) determined installed purely from binary resolution (known install paths + a where/which PATH search), with no fallback when that lookup fails for reasons unrelated to whether the CLI is actually installed (stale PATH inherited by a long-running/background process, the binary having moved, an install method not yet catalogued, etc.) — even though ~/.claude/settings.json on disk proves the tool was installed and used before. Upstream 9router's equivalent route already has this exact fallback. A new withSettingsFallback() (src/shared/services/cliInstallFallback.ts) restores 9router parity: when the binary lookup's own reason is "not_found" (never for deliberate security rejections like unsafe/relative env overrides or symlink escapes) and the tool's settings file exists on disk, installed now reports true. Regression guard: tests/unit/repro-6701-claude-detect-fallback.test.ts.

  • fix(cli): per-agent AgentBridge DNS toggle was broken for 8 of the 9 supported agents, and a failed MITM startup step could orphan the spawned proxy child — addDNSEntry/removeDNSEntry (src/mitm/dns/dnsConfig.ts) always resolved the legacy Antigravity default hosts regardless of which agent's toggle was flipped, so enabling DNS for Cursor/Codex/Claude Code/etc. silently added only daily-cloudcode-pa.googleapis.com while the DB recorded dns_enabled=true for the selected agent. Both functions now accept an optional agentId and resolve hosts via ALL_TARGETS; POST /api/tools/agent-bridge/agents/[id]/dns passes the route's id through and now returns 404 for an id that doesn't match a known target instead of silently falling back. Separately, startMitmInternal() (src/mitm/manager.ts) now wraps generateCert() (log + rethrow), the provisionDnsEntries() call, and the PID-file write in try/catch so a mid-startup failure can't orphan the already-spawned MITM child process. On Windows, addDNSEntries/removeDNSEntries also batch every missing/present entry into a single elevated PowerShell invocation instead of one UAC prompt per host line. Regression guard: tests/unit/dns-config-generic.test.ts (agent-specific resolution + batching), tests/unit/agent-bridge-dns-route-validation.test.ts (404 for unknown agent id). (#6338 — thanks @hamsa0x7)

  • fix(guardrails): Vision Bridge's individual-model auto-reroute (route an image-bearing request straight to a vision-capable model instead of describe-then-forward) could bypass a policy-restricted API key's model allowlist/budget (#6640) — VisionBridgeGuardrail.preCall() (src/lib/guardrails/visionBridge.ts) swaps body.model to the best available vision-capable model, but that swap happens in the guardrail pipeline AFTER chat.ts already called enforceApiKeyPolicy() against the ORIGINAL model, so a key scoped to a narrow allowedModels list could still execute against an unvetted (and possibly costlier) vision model the reroute picked. chat.ts now re-validates any guardrail-driven model change against the same per-key allowlist (isModelAllowedForKey) before honoring it, falling back to the original already-approved model when the reroute target is not allowed. The reroute path also now honors an explicit settings.visionBridgeModel operator override (previously ignored, unlike the combo/describe path a few lines below it, which already respects it via getVisionBridgeConfig). Regression guard: tests/unit/guardrails/visionBridge.test.ts (22 tests). (thanks @herjarsa)

  • fix(auth): an API key restricted via allowedModels/allowedCombos could bypass that restriction entirely over the Codex Responses-over-WebSocket bridge (#6564) — prepare() in src/app/api/internal/codex-responses-ws/route.ts authenticated the WS bridge's API key (authenticate()/authorizeWebSocketHandshake()) and honored allowedConnections, but never called enforceApiKeyPolicy(), the same model/combo policy gate the HTTP /v1/responses path enforces via handleChat() — so a key scoped to e.g. combo/model-1.0 could still reach a direct Codex model like gpt-5.5 through this transport, as long as an eligible Codex OAuth connection existed. The bridge's WS auth token arrives via query params (api_key/token/access_token), not a normal Authorization header, so a new enforceCodexWsApiKeyPolicy() builds an equivalent Request carrying an explicit Authorization: Bearer <apiKey> header and calls enforceApiKeyPolicy() against the CLIENT-requested model, before any Codex-specific model remapping or credential selection. Regression guard: tests/unit/codex-ws-policy-enforcement-6564.test.ts (a model-restricted key is rejected 403 before reaching credential selection; a combo-restricted key is rejected 403 requesting a disallowed combo; a key that DOES allow the requested model still proceeds past policy). (thanks @Squawk7777 for the report and an independent fix via #6565)

  • fix(security): loopback-gate /api/middleware/* so a leaked JWT over a tunnel can't install or trigger a middleware hook — middleware hooks compile + run arbitrary JS via new vm.Script on the request hot path (src/lib/middleware/registry.ts), the same RCE class as the already-gated /api/plugins/*; /api/middleware/ is now in LOCAL_ONLY_API_PREFIXES so loopback enforcement runs unconditionally before any auth check (Hard Rules #15 + #17). Regression guard: tests/unit/route-guard-middleware-local-only.test.ts. (#6541) — see PR. (thanks @developerjillur)

  • fix(startup): AgentBridge's MITM server no longer fails to start with ROUTER_API_KEY is required on a normal install (#6403) — POST /api/tools/agent-bridge/server resolved the spawned MITM child's router key from only an explicit apiKey body field (never sent by the AgentBridge UI — the schema has no such field) and the ROUTER_API_KEY env var (unset by default), so startMitm() always received "" and the child hard-exited, even though OmniRoute already had a usable API key in its own DB. A new resolveRouterApiKey() now falls back to pickApiKeyForInternalUse() (the same DB-backed selector the combo-health-check / cloud-sync internal probes use), resolving in order: explicit key → ROUTER_API_KEY env → an existing DB key. Regression guard: tests/unit/agentbridge-mitm-router-key-6403.test.ts.

  • fix(providers): deploying a Cloudflare relay Worker from Dashboard → System → Proxy pool → Cloudflare relay failed immediately with Cloudflare Worker upload failed: Content-Type must be one of: application/javascript, text/javascript, multipart/form-data, even with a valid token/account (#6416) — the Worker-script upload built a native FormData and let fetch derive the multipart Content-Type automatically, but in production globalThis.fetch is patched with node_modules/undici's own fetch (open-sse/utils/proxyFetch.ts), whose FormData/Request classes differ from the runtime's global FormData (same cross-realm class mismatch already fixed once for image edits in #3273); passing a native FormData instance through undici's patched fetch made it serialize the body as the literal string "[object FormData]" with Content-Type: text/plain;charset=UTF-8, which Cloudflare rejects outright. buildCloudflareWorkerUploadRequest() (src/lib/proxyRelay/cloudflareWorkerScript.ts) now builds the multipart body as a raw Buffer with an explicit boundary and Content-Type: multipart/form-data; boundary=… header, accepted verbatim by any fetch implementation. Regression guard: tests/unit/cloudflare-worker-upload-content-type-6416.test.ts + updated tests/unit/relay-deploy-5128.test.ts.

  • fix(security): SSRF-guard the provider-validation probes so they can no longer be used as an open relay to cloud-metadata endpoints — directHttpsRequest() (web-cookie / NVIDIA / Z.AI validation, all with a caller-controllable baseUrl) ran with guard:"none" + allowRedirect:true; it now applies getProviderValidationGuard() (default block-metadata: LAN/localhost allowed, 169.254.169.254/link-local IMDS rejected, opt-out via OMNIROUTE_ALLOW_PRIVATE_PROVIDER_URLS) and allowRedirect:false so a provider can't 3xx-redirect the probe to metadata past the initial-URL guard. Regression guard: tests/unit/provider-validation-ssrf-guard.test.ts. (#6542) — see PR. (thanks @developerjillur)

  • fix(startup): AgentBridge's MITM proxy served a mismatched cert for 3 of the 4 antigravity/cloudcode-pa hosts it terminates TLS for, breaking interception (#6494) — src/mitm/server.cjs's TARGET_HOSTS decrypts all 4 hosts locally (daily-cloudcode-pa.googleapis.com, cloudcode-pa.googleapis.com, daily-cloudcode-pa.sandbox.googleapis.com, autopush-cloudcode-pa.sandbox.googleapis.com), but src/mitm/cert/generate.ts's self-signed cert only carried a SAN entry for the first host — a request to any of the other 3 got served a cert whose CN/SAN didn't match (confirmed via curl -k https://cloudcode-pa.googleapis.com/ showing CN=daily-cloudcode-pa.googleapis.com). generateCert() now sources its host list from ANTIGRAVITY_TARGET.hosts (src/mitm/targets/antigravity.ts, the single authoritative registry already kept in lock-step with server.cjs/dnsConfig.ts/mitmToolHosts.ts by their own drift tests) and emits a SAN entry for all 4 hosts instead of hard-coding a second, incomplete copy. Regression guard: tests/unit/agentbridge-antigravity-cert-hosts-6494.test.ts (asserts the host list covers all 4 hosts and that the real generated cert's SAN includes each one).

  • fix(resilience): a priority combo never fell back when a target masked credit/quota exhaustion behind an HTTP 200 (#6427) — validateResponseQuality() (open-sse/services/combo/validateQuality.ts) only inspected the response body's top-level error field when choices was ALSO missing/empty (the narrower #3424 empty-completion case); a masked 200 that echoed a non-empty stub choices alongside a structured error object, or a known exhaustion phrase (e.g. "insufficient credits", "quota exceeded") in the error envelope, slipped through as "valid" and the combo kept returning the dead target's response forever instead of failing over. The quality check now inspects the error envelope — a top-level OpenAI-shape error object, or a bounded, case-insensitive exhaustion-phrase match against error.message/error.code/error.type/top-level message/detail — unconditionally, before any shape-specific branch, and regardless of whether choices/output also look structurally present. The check never inspects choices[].message.content, so a legitimate completion that merely mentions "quota" or "credits" in assistant prose is not misclassified. Regression guard: tests/unit/masked-200-exhaustion-fallback-6427.test.ts.

  • fix(security): fail-closed CORS for the cookie/session-authed cloud-agent management routes — getCloudAgentCorsHeaders() reflected any caller's Origin and paired it with Allow-Credentials: true (a CSRF/exfil hole); it now defers to the central allowlist (resolveAllowedOrigin), echoes only an allowlisted origin with Vary: Origin, and emits Allow-Credentials only for an explicitly allowlisted origin — never for a CORS_ALLOW_ALL wildcard echo. Regression guard: tests/unit/cloud-agent-cors-failclosed.test.ts. (#6543) — see PR. (thanks @developerjillur)

  • fix(compression): adaptive-compression ladder ranked 6 real catalog engines (ccr, ionizer, relevance, llmlingua, llm, read-lifecycle) as if they didn't exist (#6533) — ladder.ts's AGGRESSIVENESS and REDUCTION_FACTOR maps only covered the 7 engines wired into DEFAULT_LADDER (session-dedup/rtk/headroom/lite/caveman/aggressive/ultra); every other engine registered in open-sse/services/compression/engines/index.ts — including ccr/llmlingua, which the ladder doc comment already says are intentionally addable via ladderOverride — fell through to aggressivenessOf()'s ?? 0 default (same rank as "off") and expectedReductionFactor()'s generic ?? 0.9 fallback, so floor-mode escalation could not rank or escalate past them once added to a custom ladder. Both maps now carry entries for all 6 missing real engines, rescaled ×10 (off:0ultra:70) and placed by each engine's documented stackPriority (ionizer between rtk/headroom, relevance before caveman, llmlingua/llm between aggressive and ultra, etc.); mcpAccessibility — named in the report — is not a registered CompressionEngine (it's a separate MCP tool-response truncation mechanism) and was correctly left out. Regression guard: tests/unit/ladder-engine-maps-6533.test.ts (asserts every id from listCompressionEngines() ranks above "off" with a non-default reduction factor). (thanks @chirag127)

  • fix(api): tool-call arguments could render as [object Object] sequences instead of the real JSON through the /anthropic (Anthropic-shape /messages) routing path (#6459) — appendToolCallArgumentDelta() (open-sse/utils/toolCallArguments.ts), the shared accumulator the streaming openai-to-claude response translator, openai-responses translator, and responsesTransformer all call to build up a tool call's arguments/input_json_delta buffer, treated any non-string incoming fragment as an empty string. Some upstreams deliver the full tool_calls[].function.arguments value as an already-parsed JSON object/array instead of the OpenAI-contracted JSON-encoded string; the old code silently discarded that fragment, leaving tool_use.input empty, and left downstream buffers open to a plain string coercion of the object ([object Object]) once client-side concatenation kicked in. appendToolCallArgumentDelta() now JSON.stringify()s a non-string, non-null object/array fragment into a valid JSON fragment instead of dropping it, so the assembled partial_json always parses back into the original structured value. Regression guard: tests/unit/anthropic-toolcall-args-6459.test.ts. (thanks @chirag127)

  • fix(providers): fusion combo returned the opaque "All fusion panel models failed" 503 even when only a minority of panel members were actually cooling down / rate-limited, and a user-supplied fusionTuning.minPanel=1 was silently overridden (#6454) — handleFusionChat() hard-clamped the quorum floor via Math.min(Math.max(2, cfg.minPanel), panel.length), so an operator-configured minPanel=1 never took effect: collectPanel()'s straggler-grace timer only starts once ok >= minPanel, and with the floor forced to 2 a single fast success plus N slow-failing stragglers never reached quorum, so the panel sat waiting instead of degrading to the survivor. Per-member failure reasons (straggler_dropped/timeout/threw/status_XXX/empty_content/unparseable) were also logged server-side but never surfaced in the 503 body, leaving operators unable to tell a rate-limit fan-fail from a broader outage. Fixed by honoring Math.max(1, cfg.minPanel) and threading a failures: Array<{ model, reason }> collector into the 503 message (model=reason per entry) — production fix already merged via #6521; this entry backfills the missing CHANGELOG bullet and adds an 11-member, fusion-free-scale regression test matching the original repro shape (a cooling minority must not sink a healthy majority; a genuinely all-failed panel still returns the documented 503). Regression guard: tests/unit/services/fusion-min-panel-and-failure-detail.test.ts + tests/unit/fusion-partial-panel-failure-6454.test.ts. (thanks @chirag127)

  • fix(providers): fusion combo strategy silently returned a panel member's raw answer instead of the configured config.judgeModel synthesis (#6455) — handleFusionChat()'s single-survivor "degrade gracefully" path (added for #6454) returned the lone panel answer directly whenever only one panelist succeeded, regardless of whether an explicit judgeModel was configured; with the default minPanel: 2 and a 2-model panel, any single flaky/rate-limited panelist forced this path on every request, so the configured judge (e.g. auto/claude-opus) was never invoked and the client-visible .model reflected whichever panelist happened to survive. The judge is now still invoked to synthesize a lone surviving answer whenever judgeModel is explicitly configured; the cheap direct-answer shortcut is kept only for the implicit case (no judgeModel set, where the "judge" is just panel[0]). Regression guard: tests/unit/fusion-judge-model-6455.test.ts + updated tests/unit/combo-fusion-strategy.test.ts. (thanks @chirag127)

  • feat(combo): sanitized diagnostic trace on an auto-combo terminal failure — instead of an opaque 503, a terminal combo failure now returns a whitelist-projected trace (candidate pool size, attempted count, excluded provider/reason codes, attempt order, and a terminal-reason code) via the new errorResponseWithComboDiagnostics()/sanitizeComboDiagnostics() in open-sse/utils/error.ts — provider/model ids and enumerated reason codes only, never keys/tokens/bodies, length- and count-capped. A reasoning-budget-exhausted panel now returns an actionable "increase max_tokens" message rather than a blind retry-limit 503. Regression guard: tests/unit/combo-diagnostics-trace.test.ts. (#6545) — see PR. (thanks @developerjillur)

  • fix(providers): image/diffusion models discovered from an upstream catalog (e.g. HuggingFace's live /v1/models) are no longer advertised as chat models (#6457) — the chat catalog builder defaulted synced models with no modality info to endpoints: ["chat"], so huggingface/stabilityai/stable-diffusion-xl-base-1.0 showed up in the chat /v1/models listing and returned 400 "not a chat model" when called. catalog.ts now skips any synced model already registered as an image model for that provider (via the new isRegisteredImageModel()), leaving getAllImageModels() to list it with the correct type: "image". Regression guard: tests/unit/image-model-not-in-chat-catalog-6457.test.ts.

  • fix(resilience): combo session stickiness never released a pin on a credits-exhausted/banned/expired account, permanently defeating failover for that conversation (#6692) — applySessionStickiness() (open-sse/services/combo/sessionStickiness.ts) gated the sticky pin only on 5h/weekly usage-percentage headroom, which is orthogonal to account availability: a credits_exhausted/banned/expired connection, or one still inside its rateLimitedUntil cooldown, reports perfectly healthy headroom, so the pin was force-promoted back to the front of the target list on every subsequent turn. clearStickyBinding() also had zero call sites in combo.ts's failure paths, so a quality-validation-rejected 200 (a masked daily-cap refusal) never released the pin either. The gate now also resolves the bound connection's terminal status/cooldown via a new injectable fetcher seam (fail-open on lookup errors, mirroring the existing saturation fetcher), and combo.ts's two dispatchers (handleComboChat/handleRoundRobinCombo) release the pin immediately at both their connection-exhaustion classification point and their quality-validation-failure branch via the new releaseStickyPinOnFailure(). Regression guard: tests/unit/repro-6692-sticky-terminal.test.ts + extended tests/unit/combo-session-stickiness.test.ts.

  • fix(test): replace the bare expect(true).toBe(true) tautology in playground-api-tab.test.tsx's SSE test and close the check:test-masking gap that let it slip through for a full cycle (#6404) — a prior pass (#6548) had already swapped the literal to expect(sendBtn).toBeDefined(), but that stayed just as vacuous: the test's fetch mock returned an empty /v1/models list, so ApiTab's Send button is always disabled (!selectedModel) and the SSE branch never runs — the "SSE infra is verified" comment was never true. The test now mocks a real model, drives the model <select> to enable Send, asserts sendBtn.disabled === false before clicking, and asserts the streamed SSE delta ("Hello!") actually reached the response editor. Root cause on the detector side: check-test-masking.mjs's tautology subcheck only compares base-vs-HEAD counts within a PR's own diff (headExtTaut > baseExtTaut) and no-ops locally when GITHUB_BASE_SHA/GITHUB_BASE_REF are unset ("sem base ref — pulando") — so a tautology merged once, or checked with a bare local run, was invisible forever after. Added a new always-on, PR-independent absolute-floor scan (scanBareTautologies + countBareTautologies) over every git-tracked test file for the bare expect(true).toBe(true) / assert.equal(1,1) / assert.strictEqual(1,1) patterns specifically (deliberately excluding assert.ok(true), which has ~15 pre-existing verified-legitimate try/catch-fallback uses repo-wide and stays governed by the lenient diff-only subcheck) — verified zero pre-existing hits repo-wide once this file was fixed, so the new floor is safe to enforce unconditionally. Regression guard: tests/unit/check-test-masking.test.ts (new scanBareTautologies/countBareTautologies cases) + tests/unit/ui/playground-api-tab.test.tsx. (thanks @chirag127)

  • fix(oauth): Codex/ChatGPT (and every other OAuth provider) connection stays stuck showing "Auth Failed" even after a genuinely successful token refresh (#6352) — updateProviderCredentials() (the shared onPersist callback for the manual "Refresh token" route, the reactive per-request refresh in chat.ts, and the Codex/Claude auth-file importers) correctly reused the stored refresh_token, persisted the new access_token, and replaced a rotated refresh_token, but never cleared the stale testStatus/lastError*/errorCode fields left over from a prior expired/invalid refresh or upstream 401/403 — only the separate background health-check sweep did that clearing. A successful refresh now resets testStatus to "active" and clears lastError, lastErrorAt, lastErrorType, lastErrorSource, and errorCode (an explicit testStatus from the caller still wins). Regression guard: tests/unit/codex-oauth-refresh-persist-6352.test.ts.

  • perf(health): short-TTL (1s) cache for the frequently-polled GET /api/monitoring/health payload — rebuilding it every request (DB reads + status aggregation across 8 subsystems) was wasteful under rapid polling; the cache stays near-real-time and is invalidated immediately on DELETE (circuit-breaker reset) so a manual reset is reflected at once. Regression guard: tests/integration/monitoring-health-cache.test.ts. (#6553) — see PR. (thanks @developerjillur)

  • fix(resilience): headroom combo routing did not always select the Codex account with the most free quota (#6379) — orderTargetsByHeadroom (open-sse/services/combo/quotaStrategies.ts) already loaded the per-connection DB snapshot (with decrypted credentials) via expandTargetsByQuotaAwareConnections, but discarded it before calling getSaturation; for Codex, fetchCodexSaturation forwards straight to fetchCodexQuota(connectionId, connection), which needs connection (or a prior registerCodexConnection() call, which never happens before headroom ranking runs) to read accessToken — so it returned null for every candidate, saturation failed open to 0 across the board, and ranking fell back to the original combo order regardless of actual free quota. getSaturation() and the headroom SaturationFetcher seam now accept and thread the loaded connection snapshot through to fetchCodexQuota. Kilo's dup flag vs #5903 was a false positive — that issue is about session-sticky reset-aware/least-used selection, not headroom's Codex saturation lookup. Regression guard: tests/unit/headroom-codex-quota-snapshot-6379.test.ts. (thanks @eidoog)

  • fix(providers): custom models a provider actually has are no longer dropped from the Free Provider Rankings when both "Configured only" and "Available only" filters are applied (#6368), follow-up to #6150 — freeProviderRankings.ts::getProviderModels() only ever walked the static open-sse/config/providerRegistry.ts catalog, so a user-added custom model (e.g. a Puter claude-fable-5 model saved as "Claude Fable 5") never entered the candidate model list the ranking scores against, and could never survive the #6150 configured/available filters even when actually configured and available. It now additively merges the provider's custom models (db/models.ts::getCustomModels) into that candidate list via a new pure, de-duping mergeProviderModels() helper, before scoring/filtering runs — catalog free/paid filtering elsewhere is untouched. Regression guard: tests/unit/free-provider-rankings-custom-models-6368.test.ts. (thanks @shabeer)

  • fix(playground): accept a valid dashboard session for GET/POST /api/playground/presets under REQUIRE_API_KEY=true — the Playground page calls this route with a cookie/session and no API key, which previously 401'd the authenticated dashboard; checkAuth() now accepts a management/dashboard session (requireManagementAuth) as an alternative to an API key, while a presented API key must still be valid and the anonymous-allowed default is preserved. Regression guard: tests/integration/presets-dashboard-auth.test.ts. (#6554) — see PR. (thanks @developerjillur)

  • fix(providers): cloudflare-ai no longer silently drops image/non-text content parts (#6390) — transformRequest()'s flattenContent() (added for #2539 to satisfy the Workers AI /ai/v1/chat/completions plain-string content requirement) mapped any non-text OpenAI content part (e.g. image_url) to "" and joined the rest, so a request carrying an image quietly went out as text-only with the attachment gone and no error surfaced. It now throws a clear error on the first non-text part instead of dropping it silently, which the existing top-level chatCore.ts catch already routes through buildErrorBody()/sanitizeErrorMessage() (same pattern as buildUrl()'s missing-Account-ID error). Regression guard: tests/unit/cloudflare-ai-image-parts-6390.test.ts.

  • fix(providers): stop Antigravity connections from falsely reporting all-accounts quota-exhausted (#6295) — genericQuotaFetcher.ts::percentUsedForQuota() ignored the fractionReported flag and defaulted an unreported model's remainingPercentage to 0, which computed as 100% used; since convertUsageToQuotaInfo() takes the worst-case window across a connection, a single model with no reported fraction dragged the whole account into limitReached and quotaPreflight skipped it. percentUsedForQuota() now returns null (unknown, window ignored) whenever fractionReported === false, before falling back to remainingPercentage. Regression guard: tests/unit/generic-quota-fetcher.test.ts.

  • fix(fusion): the fusion judge no longer replays a panel member's answer via an idempotency-key collision — fusion's panel + judge sub-requests re-enter chatCore sharing the client's headers, so they derived the same Idempotency-Key/x-request-id and a panel answer saved under the key was replayed by the judge's check ~1ms later (inside the 5s window), returning a panel member's answer instead of the judge synthesis (observed live on nexa/conversation-fusion). composeIdempotencyKey() now namespaces the key by target provider/model + a digest of the request messages, so sub-requests can't collide while a genuine client retry (same key/model/body) still replays. Regression guard: tests/unit/idempotency-fusion-collision.test.ts. (#6558) — see PR. (thanks @developerjillur)

  • fix(providers): grok-cli (Grok Build) now strips reasoning_effort/reasoning before forwarding the request (#6288) — Claude Code sends reasoning_effort on every request (routing the Opus slot), which Grok Build's upstream chat-proxy endpoint rejects with a 400; transformRequest()'s existing UNSUPPORTED sampling-param strip list (#5273) never covered it. Regression guard: tests/unit/grok-cli-reasoning-strip-6288.test.ts.

  • fix(cli): omniroute serve no longer hangs silently on a readiness timeout (#6321) — the child server's stdout was piped to "ignore" whenever --log/OMNIROUTE_SHOW_LOG wasn't set (the default), discarding any debug output, and runWithSupervisor's waitForServer(...).then((up) => { if (up) {...} }) had no else branch, so a boot that never became ready produced zero further output after "⏳ Starting server...". Stdout is now buffered alongside stderr (ServerSupervisor.getRecentLog()), and a timeout prints a clear diagnostic plus the buffered output instead of staying silent. Does not by itself explain why boot never completes on a given machine — see the issue for further reproduction. Regression guard: tests/unit/cli-serve-readiness-timeout-6321.test.ts.

  • fix(pricing): Pricing Sync dashboard no longer stuck on "Next Sync: Never" / "Synced Models: 0" (#6325) — pricingSync.ts kept sync state (lastSyncTime, lastSyncModelCount) in module-level vars, but the background periodic sync (instrumentation-node.ts) and the dashboard status route (/api/pricing/sync) each import the module from separate Next.js standalone webpack chunks, giving each its own independent state; getSyncStatus() read the (empty) API-route instance's vars. Sync status is now additionally persisted to a new pricing_sync_status key_value namespace and getSyncStatus() falls back to it when the local module instance never ran a sync itself. Regression guard: tests/unit/pricing-sync-cross-instance.test.ts.

  • fix(api): stop spuriously 403-ing "Invalid request origin" on POST /api/providers/health-autopilot/actions for Docker/LAN dashboard requests (#6277) — the route carried a duplicate per-route validateBrowserMutationOrigin check re-added by the v3.8.42 release squash after PR #5278 centralized origin enforcement in the authz pipeline; the pipeline strips PEER_IP_HEADER before forwarding, so the stale duplicate check could no longer resolve the LAN "direct-local-host" candidate and rejected legitimate same-origin LAN mutations (e.g. clicking "remove cooldown" when accessed via a LAN IP). Removed the duplicate check — origin validation is now solely enforced by the centralized pipeline check, which already handles this case correctly. Regression guard: tests/unit/serial/provider-health-autopilot.test.ts.

  • fix(resilience): a bare, unrecognized 403 from a no-credential (authType:"none") provider like mimocode or theoldllm no longer permanently bans the connection (#6315, #6345) — classifyProviderError()'s 403 branch only exempted apikey providers from the terminal FORBIDDEN classification, so these free/stateless proxies (no real account/credential to revoke) fell through to FORBIDDEN on the first unmatched 403 and got isActive:false, testStatus:"banned" with no cooldown or retry. The exemption now also covers authType:"none" providers, returning null (recoverable) so the existing connection-cooldown/retry layer handles it. Regression guard: tests/unit/errorClassifier-noauth-403-6315.test.ts.

  • fix(providers): the Auggie (Augment CLI) executor no longer fails on Windows with spawn EINVAL (#6304) — the global-npm install exposes auggie as a .cmd shim, which Node's child_process.spawn cannot launch on win32 without shell: true. Both spawn sites (streaming + the auggie --version test) now go through a shared buildAuggieSpawnOptions() that sets shell: process.platform === "win32"; the argv (built by buildAuggieArgs() with a registry-validated model and a trailing -- end-of-options marker) is unchanged, so the argument-injection surface stays closed on non-Windows. Regression guard: tests/unit/auggie-win32-spawn-6304.test.ts.

  • fix(api): the dashboard "Test model" action is now a clean connection test (#6240) — modelTestRunner sent its probe request without an explicit compression override, so whenever the operator's global compression.enabled flag was on the test call inherited compression (and any Output-Styles system prompt), polluting the result. The internal test requests now send X-OmniRoute-Compression: off, and chatCore honors an explicit off header even when compression.enabled is globally true. Regression guards: tests/unit/model-test-runner-compression-off-6240.test.ts, tests/integration/test-model-compression-off-6240.test.ts.

  • fix(startup): an update/restart could crash the whole server at boot with TypeError: Cannot create property 'message' on string 'Database closed', masking the real failure and 500-ing every request until manually restarted (#6560, plausibly the root cause of #6594's post-upgrade 500) — driverFactory.ts::preInitSqlJs() cached its sql.js WASM adapter per file path in a globalThis-backed map for idempotency, but never checked whether the cached adapter had since been closed (e.g. by gracefulShutdown/resetDbInstance racing a reload); reusing that dead handle made the very next query throw sql.js's own bare string "Database closed" (not an Error) straight out of instrumentation-node.ts's previously-unguarded ensureDbInitialized() call, and Next.js's internal registerInstrumentation() wrapper unconditionally does err.message = ... on whatever register() rejects with — assigning .message on a primitive string throws in strict mode, so the secondary TypeError is what actually crashed the process. Fixed in two parts: preInitSqlJs() now evicts a closed cached adapter and creates a fresh one instead of returning it; a new ensureDbReadyForBoot() wraps the DB-init call, normalizes any non-Error throw via normalizeBootError(), and retries once specifically for a transient "database closed" message (now succeeding against the fresh adapter) before re-throwing anything else as a real Error. Regression guard: tests/unit/instrumentation-database-closed-6560.test.ts.

  • fix(api): POST /api/keys no longer hangs 20–90+ seconds on a fresh install, even with valid auth (#6570) — cloudEnabled defaults to true in src/lib/db/settings.ts::getSettings() on any install with no persisted settings row (i.e. every fresh install), so the create-key handler's unconditional await syncKeysToCloudIfEnabled() always attempted a real outbound fetch() to CLOUD_URL via syncToCloud(); when that endpoint is unset/unreachable/slow, the HTTP response blocked until the request settled or timed out — unlike sibling routes (POST /api/keys/:id/regenerate, GET /api/combos), which never touch this side effect at all. src/app/api/keys/route.ts now dispatches syncKeysToCloudIfEnabled() fire-and-forget (void) instead of awaiting it; its internal try/catch already logs failures, so cloud sync still runs, it just no longer blocks the response. Regression guard: tests/unit/api-keys-create-no-hang-6570.test.ts (asserts the route resolves in well under 2s even when the Cloud-sync fetch() is stubbed to never settle) + updated tests/integration/api-keys.test.ts (the pre-existing "triggers cloud sync"/"still succeeds when cloud sync fails" tests, which previously hung indefinitely on this exact path, now await the fire-and-forget sync tick before asserting).

  • fix(api): editing any existing OpenAI Codex provider connection in the dashboard returned "Invalid request" and the edit could never be saved (#6562) — createProviderConnection() (src/lib/db/providers.ts) auto-increments a new connection's priority to MAX(priority)+1 per provider with no upper bound, and OAuth-imported connections (Codex codex-auth/import / import-bulk, up to 50 accounts per call, callable repeatedly — the standard Codex bulk-account-rotation workflow) never pass through createProviderSchema's Zod validation at all, so nothing ever capped that value; EditConnectionModal's handleSubmit always resends the connection's current priority unchanged on every save, and updateProviderConnectionSchema capped priority/globalPriority at max(100) — a UI-only ceiling the create path never enforced — so the first edit of any connection whose priority had already grown past 100 (routine once a Codex account count exceeds 100) failed validation regardless of which field the user actually changed. Raised the ceiling to max(100_000) on both fields — still bounded (a genuinely out-of-range value is still rejected), just wide enough to accept priorities the app itself already produces. Regression guard: tests/unit/codex-connection-edit-6562.test.ts (a Codex OAuth connection whose priority already exceeds the old 100 cap now validates + persists on edit; a control payload with a still-genuinely-invalid priority is still rejected with "Invalid request").

  • mimocode: rotate accounts on MiMoCode's rate-limit-style 400s (body-classified) instead of failing on the first account; malformed 400s still fail fast with the real upstream error (#6648 — thanks @pizzav-xyz)

  • fix(cli): compression CLI REST fallback now reads/writes the canonical defaultMode field (surfaced as strategy) instead of a nonexistent engine key, and table output renders nested objects as JSON instead of [object Object] (#6571 — thanks @charleszolot)

  • fix(providers): web-cookie providers without a providerRegistry.ts entry (lmarena, gemini-business, poe-web, venice-web, v0-vercel-web) now report unsupported: true instead of silently "OK" (#6309) — validateWebCookieProvider() (src/lib/providers/validation.ts) previously required a registry entry and returned "Provider not found in registry" for these; a fallback to WEB_COOKIE_PROVIDERS[provider].website was proposed, but live verification showed the ${website}/models probe does not reliably signal session validity for these providers (redirects/SPA 200s regardless of cookie validity — e.g. lmarena's real API is arena.ai, not lmarena.ai; Poe's real endpoint is a GraphQL POST, not a REST /models), so it would report an expired or garbage cookie as valid. Until each provider has a verified, side-effect-free auth probe against its real API host, the fallback now returns unsupported (no network call) instead of a false positive. Regression guard: tests/unit/web-cookie-validation-fallback.test.ts. (thanks @oyi77)

  • fix(api): POST /api/middleware/hooks and PUT /api/middleware/hooks/[name] no longer leak raw internal error messages in their 500 responses (#6645 — thanks @chirag127) — both catch blocks returned error?.message directly (Hard Rule #12), which could surface internal SQLite path fragments on a DB failure; both now route through sanitizeErrorMessage() from open-sse/utils/error.ts. Regression guard: tests/unit/middleware-hooks-error-sanitization.test.ts.

  • fix(docker): compile better-sqlite3 for the server Docker image (Dokploy/self-hosted builds) via a direct node-gyp rebuild inside node_modules/better-sqlite3, instead of npm rebuild better-sqlite3 (#6700) — the builder stage installs dependencies with npm ci --ignore-scripts (deliberate: closes the supply-chain surface where a transitive dep's install script runs arbitrary code) and re-enables the native build for the one package that needs it; npm rebuild <pkg> re-runs that indirectly through the package's own install script, which under npm 11 depends on npm's script-allowlist machinery correctly re-enabling it — some self-hosted build environments (e.g. Dokploy) hit a broken/mismatched native binding through that indirection. Invoking node-gyp rebuild directly bypasses npm's script-running layer entirely and is deterministic regardless of npm version. Regression guard: tests/unit/dockerfile-better-sqlite3-node-gyp-6700.test.ts. (thanks @nowhats-br)

  • fix(providers): the Cloudflare relay Worker deploy fix in #6416/#6618 still failed uploads in practice — it changed the multipart Content-Type but kept the emitted worker source as an ES module (export default { fetch(...) }) with main_module metadata; Cloudflare's Workers upload API parses a plain application/javascript script part as Service Worker syntax regardless of the main_module metadata field, and main_module requires the script to actually be an ES module (top-level export), so the mismatch still rejected the upload (#6496). buildCloudflareWorkerScript() (src/lib/proxyRelay/cloudflareWorkerScript.ts) now emits Service Worker syntax (addEventListener("fetch", ...), no top-level export) and the upload metadata uses body_part instead of main_module. Regression guard: tests/unit/relay-deploy-5128.test.ts (asserts the emitted script has no export default, registers a fetch listener, and the upload metadata carries body_part/omits main_module; also proves the inlined isPrivateHostname() SSRF guard still rejects bracketed IPv6 loopback/ULA hosts like [::1]/[fd00::1] after the script-body rewrite). (thanks @SeaXen)

  • fix(providers): ChatGPT Web (chatgpt-web) responses rendered raw ChatGPT UI citation markup — private-use marker tokens (e.g. citeturn0search0) and url… inline-link markers — instead of real Markdown links, since these only ever get resolved client-side by chatgpt.com's own JS using message.metadata.content_references (#6635) — cleanChatGptText() now resolves content_references (grouped webpages, footnote sources, inline webpage/url mentions) into [label](url) Markdown links for both the streaming and non-streaming response builders, and for the GPT-5.5 Pro stream_handoff polled-answer path, falling back to stripping any marker that has no resolvable source instead of leaking the raw private-use bytes. The citation parsing/rendering logic was extracted into a new pure sibling module (open-sse/executors/chatgpt-web/citations.ts) to keep the executor under the frozen file-size cap. Regression guard: tests/unit/chatgpt-web-citations.test.ts (non-streaming citation resolution, streaming marker buffering across split SSE chunks, and the Pro-handoff polled-answer path). (thanks @Thinkscape)

  • fix(dashboard): the live-dashboard WebSocket descriptor handshake (GET /api/v1/ws?handshake=1) and the lightweight GET /api/health/ping liveness probe both 401'd for unauthenticated callers, even though both are metadata-only reads intended to be public (#6335) — clientApiPolicy required a bearer/dashboard-session before the WS route handler could even return its own wsAuth/protocol descriptor, and /api/health/ping was never added to PUBLIC_READONLY_API_ROUTE_PREFIXES despite its own docstring documenting it as "No auth required." clientApiPolicy.evaluate() now allows an anonymous {kind:"anonymous", id:"ws-handshake"} subject for GET/HEAD/OPTIONS on /api/v1/ws?handshake=1 (the route handler still performs its own real wsAuth/dashboard/API-key decision before opening the socket), and /api/health/ping is now in PUBLIC_READONLY_API_ROUTE_PREFIXES. Regression guard: tests/unit/authz/client-api-policy.test.ts (WS handshake allowed, including relative request URLs), tests/unit/public-api-routes.test.ts, and tests/unit/authz/classify.test.ts (/api/health/ping classified PUBLIC). (thanks @JxnLexn)

  • fix(providers): wire the Devin cloud-agent provider into the generic provider-page validator and static model catalog, matching the existing jules cloud-agent pattern (#6142)

  • fix(providers): honor a provider-level proxy assigned to no-auth providers like MiMoCode Free (#6272)

  • fix(api): merge tool_call continuation deltas that carry only id (no index) so tool-call arguments are no longer split/lost in request/response logs (#6276)

  • fix(providers): modernize the lmarena provider for the Arena.ai rebrand — route chat through arena.ai create-evaluation with Chrome TLS impersonation, seed a static Direct-chat Text/Search + Image catalog, and keep the lmarena/lma wire id for back-compat (#6280) — thanks @backryun

  • fix(providers): web-provider model discovery updated — qwen-web uses the slash-terminated models endpoint (avoiding a blocked 307 redirect), and kimi-web matches the current request shape (POST with bearer + kimi-auth cookie replay) with its catalog refreshed to the current non-agent models (#6308 — thanks @janeza2).

  • fix(logs): the request-log detail modal no longer reopens by itself after being closed — a stale in-flight detail refresh resolved after close and re-triggered the modal open state (#6323 — thanks @xz-dev).

  • fix(providers): update SenseNova Token Plan support — register the token-plan model ids/constants and adjust the SenseNova registry so token-plan accounts route correctly (#6330 — thanks @xz-dev).

  • fix(providers): give v0-vercel-web its own alias so its credentials are detected (#6343)

  • fix(providers): route AgentRouter key validation through the CC wire image so a valid key no longer 403s as "Invalid API key" (#6377)

  • fix(db): stop legacy log-archive migration from deleting the live app-logger directory and crashing startup on a stat/stream race (#6401, #6799)

  • fix(docs): document Turbopack build memory tradeoff and OMNIROUTE_USE_TURBOPACK=0 webpack fallback for RAM-constrained machines (#6409)

  • fix(compression): surface silently-dropped stacked-pipeline steps (session-dedup, ccr) and stop the aggregate inflation guard from misfiring on a genuine no-op (#6479, #6480, #6491)

  • fix(providers): honor the max_token capability override in the reasoning-token-buffer output cap (#6524)

  • fix(dashboard): the onboarding tier-flow diagram rendered broken — its SVGs lived in the repo-root images/ (not a served path); moved to public/images/ so Next.js serves them (#6538 — thanks @ianriizky).

  • fix(routing): the auto combo's no-auth candidate pool now honors a disabled provider connection's own isActive=false (the toggle on the main Providers grid card), not just the separate global blockedProviders setting — disabling opencode/mimocode/etc. via the grid toggle no longer leaves it in rotation (#6557).

  • fix(sse): server-tool literal names (e.g. web_search) are preserved in message history and tool_choice instead of being namespaced/rewritten, so follow-up turns referencing those tools keep working (#6586 — thanks @MikeTuev).

  • fix(api): recognize OpenRouter reasoning/reasoning_details in non-streaming OpenAI-to-Claude conversion (#6623)

  • fix(db): share one in-flight sql.js load across concurrent preInitSqlJs() callers to stop the boot-time thundering-herd re-decode of the whole database file (#6628)

  • fix(db): unwrap lone named-parameter objects before sql.js stmt.bind() so @/:/$-style named placeholders bind correctly instead of throwing "Wrong API use" (#6802)

  • fix(db): break probe-failed/restore loop on large storage.sqlite (#6632 — thanks @KooshaPari).

  • fix(ci): exclude check-test-masking.test.ts's own tautology fixtures from the diff-based test-masking gate and recognize in validate-release-green's failure-line detector (#6634)

  • fix(routing): recognize Kimi-style "exceeded model token limit" 400 as context overflow so combo fallback continues to the next target (#6637)

  • fix(cli): Claude Code installed via WinGet is now detected on Windows (the WinGet install path was missing from the binary lookup) (#6647 — thanks @enjoyer-hub).

  • fix(providers): removed obsolete/defunct providers from the catalog (glhf, kluster, cablyai, inclusionai) (#6675 — thanks @backryun).

  • fix(sse): requests rejected before handleChatCore (circuit-breaker/cooldown gate or combo with all targets exhausted) are now recorded in usage_history too, so a key whose traffic was entirely gate-rejected no longer shows "zero requests" in the per-API-key usage counter (#6698).

  • fix(sse): unwrap bare {function:{…}} tools so OpenAI-shape clients no longer have tools silently dropped in Claude translation. (thanks @samir-abis)

  • fix(oauth): stop merging distinct Codex OAuth logins that share an email but lack a verifiable account id, preventing silent token overwrite. (thanks @lucasjustinudin)

  • fix(codex): detect "model at capacity"/overloaded errors embedded in a 200-OK SSE stream and surface them as a real error so account fallback rotates, instead of passing them through as a successful response. (thanks @ryanngit)

  • fix(volcengine): clamp max_tokens to the VolcEngine Ark endpoint cap for the Kimi model so oversized values no longer 400. (thanks @whale9820)

  • fix(antigravity): surface aborted/malformed Gemini tool calls (e.g. MALFORMED_FUNCTION_CALL) as an explicit non-end_turn finish reason instead of a silent clean completion. (thanks @anhdiepmmk)

  • fix(routing): the reasoning-token headroom buffer clamps to the model's explicit output cap instead of inflating past it, and getExplicitModelOutputCap falls through to the registry/spec cap when a synced capability row exists without a numeric limit_output (#6714) — thanks @xz-dev

  • fix(api): omniroute health (and health components/health watch) returned Error: HTTP 404 (#6677) — bin/cli/commands/health.mjs called apiFetch("/api/health", ...), a route that was moved to GET /api/monitoring/health (src/app/api/monitoring/health/route.ts) without updating the CLI; src/app/api/health/ on disk only has degradation/route.ts and ping/route.ts, no top-level handler. runHealthCommand()/runHealthComponentsCommand() now call /api/monitoring/health and read its actual payload shape (activeConnections, circuitBreakers: {open, halfOpen, closed}, memoryUsage) instead of the old, nonexistent requests/breakers/cache/memory fields. Regression guard: tests/unit/cli-health-monitoring-route.test.ts.

  • fix(startup): webpack build broke on case-insensitive filesystems (macOS APFS default, Windows) with a casing-collision warning plus "not exported" errors in StudioConfigPane.tsx/ChatTab.tsx (#6584) — src/app/(dashboard)/dashboard/playground/components/ReasoningControls.tsx (the component) and reasoningControls.ts (the utils module) shared the same lower-cased stem in the same directory, and two importers used the extensionless form from "./reasoningControls", the exact resolution path that becomes ambiguous once casing is folded. Renamed the utils module to reasoningControlUtils.ts (no collision) and updated the 3 import sites. Regression guard: tests/unit/case-collision-6584.test.ts (scans src//open-sse/ for any same-directory, case-only filename collision). (#6584)

  • fix(build): Turbopack production build emitted an "Overly broad patterns can lead to build performance issues" warning per entry point importing src/lib/agentSkills/generator.ts (603 warnings reported on v3.8.46, up from 379 on v3.8.45) (#6582) — generator.ts's outputBase is built as path.isAbsolute(outputDir) ? outputDir : path.join(process.cwd(), outputDir), where outputDir is a runtime function parameter, not a compile-time literal, so Turbopack's build-time file-tracing analyzer can't statically narrow the several dynamic readdirSync/rmSync/readFileSync/writeFileSync call sites a few lines below and falls back to a project-wide glob; #6366's commit message claimed to "anchor the base path with a literal" but the shipped code never did. Since this fs access is legitimate and bounded (skills/<id>/SKILL.md, ~48 known IDs), next.config.mjs's turbopack.ignoreIssue (Next.js 16.2+) now suppresses this specific, known-benign diagnostic, mirroring the existing webpack.ignoreWarnings/isNextIntlExtractorDynamicImportWarning precedent already in the same file for the webpack path. Regression guard: tests/unit/next-config.test.ts (asserts the turbopack.ignoreIssue rule shape targeting src/lib/agentSkills/**).

  • fix(providers): Codex Desktop requests to gpt-5.3-codex-spark failed with [400]: Tool 'image_generation' is not supported with gpt-5.3-codex-spark, even on paid-plan accounts (#6651) — CodexExecutor.transformRequest (open-sse/executors/codex.ts) only dropped the Codex Desktop-injected image_generation hosted tool when isCodexFreePlan() matched the account's plan, with no awareness that Spark-scope Codex models reject image_generation upstream regardless of plan. dropImageGeneration now also drops it when getCodexModelScope(model) === "spark" (the existing Spark classifier from open-sse/config/codexQuotaScopes.ts), independent of account plan. Regression guard: tests/unit/codex-spark-image-generation.test.ts (thanks @alltomatos for independently catching and fixing it via #6819).

  • fix(providers): the provider quota card's weekly/session bars re-sorted by remaining percentage instead of staying in a fixed, deterministic order (#6687) — QuotaCardExpanded.tsx's sortQuotasByRemaining() (added in #5977) was applied unconditionally via useMemo(() => sortQuotasByRemaining(quotas), [quotas]), undoing the deterministic CODEX_QUOTA_ORDER/GLM_QUOTA_ORDER window order quotaParsing.ts's sortCodexOrder()/sortGlmOrder() (added in #6336) already established for Codex and the GLM family — since #6336 never touched QuotaCardExpanded.tsx, the two orderings never composed, so e.g. a Codex session window with less headroom than weekly rendered after it instead of staying first. A new hasFixedQuotaOrder() (quotaParsing.ts) and resolveQuotaDisplayOrder() (QuotaCardExpanded.tsx) now skip the remaining-% re-sort for providers with a fixed window order, threading providerId from QuotaCard.tsx through to the display layer; every other provider still gets the remaining-% sort. Regression guard: tests/unit/quota-card-expanded-fixed-order-6687.test.ts.

  • fix(i18n): pt-BR was missing 194 UI keys present in en.json — a real, silent data-sync gap, not covered by any duplicate/mislabeled #6694 (that issue's 9 providers.* keys are disjoint, present-but-untranslated sentinels caused by a separate providerText() fallback bug) (#6695) — scripts/i18n/sync-ui-keys.mjs (which mirrors newly-added en.json keys into every locale) wasn't re-run after recent en.json additions, and the CI i18n:check-ui-coverage gate only fails a locale below an 80% threshold, so pt-BR stayed green at 93.8% coverage despite the gap. Backfilled all 194 missing keys into src/i18n/messages/pt-BR.json (translated to Brazilian Portuguese, no leftover __MISSING__ markers) via npm run i18n:sync-ui -- --locale=pt-BR + manual translation. Regression guard: tests/unit/i18n-pt-br.test.ts (new case asserting full en.jsonpt-BR.json key parity, so a future drift fails a fast unit test instead of silently degrading the coverage percentage).

  • fix(startup): omniroute --mcp crashed at Node ESM link time with ERR_MODULE_NOT_FOUND for ioredis on installs where the published MCP bundle didn't happen to have ioredis rescued from a parent node_modules (#6559) — src/shared/utils/rateLimiter.ts had a top-level static import Redis from "ioredis"; that module is only ever reached via a lazy await import(...) several call-sites deep in the MCP tool chain, but esbuild's --packages=external bundling of the MCP server (scripts/build/prepublish.ts Step 8.5) still hoisted rateLimiter.ts's own static import into a real top-level ESM import in the compiled dist/open-sse/mcp-server/server.js, forcing Node to resolve ioredis at module-link time — before any --mcp startup code runs — and ioredis is not guaranteed to ship in the MCP-only bundle's node_modules. getRedisClient() now lazily imports ioredis on first use (matching the established soft-dependency pattern in src/lib/quota/redisQuotaStore.ts) while still throwing synchronously when Redis isn't configured. Regression guard: tests/unit/build/mcp-bundle-no-eager-ioredis.test.ts (bundles the real MCP server entrypoint with the exact publish-time esbuild flags and asserts no top-level static ioredis import remains, while the pre-existing lazy await import("ioredis") in redisQuotaStore.ts stays intact).

  • fix(providers): Kiro sent the adaptive-thinking additionalModelRequestFields envelope for claude-sonnet-4.5/claude-haiku-4.5, which Kiro/CodeWhisperer rejects upstream with a raw [400]: additionalModelRequestFields is not supported for this model (#6576) — buildKiroPayload() (open-sse/translator/request/openai-to-kiro.ts) gated the field on the generic Anthropic-API supportsReasoning() capability flag, which is true for both models on Anthropic's direct API but does not reflect what Kiro's CodeWhisperer backend actually accepts; only claude-sonnet-5 is confirmed adaptive-thinking-capable there. A new Kiro-specific allowlist (supportsKiroAdaptiveThinking() in open-sse/translator/request/openai-to-kiro/adaptiveThinking.ts) now gates the envelope instead. Regression guard: tests/unit/repro-6576-kiro-thinking-unsupported-model.test.ts.

  • fix(translator): Cursor's local Subagent tool call is no longer rejected with cloud_base_branch may only be specified when environment equals cloud — the Responses→Chat tool-arg cleanup (stripEmptyOptionalToolArgs) was scoped to Claude Code's Read tool only, so Cursor's Subagent tool passed through with the cloud-only cloud_base_branch: "" (Cursor treats an empty string as "specified" and rejects the call before starting the local subagent). The cleanup now covers an allowlist of Read + Subagent; arbitrary tools are still left untouched (empty strings/arrays can be valid payloads for them). Regression guard: tests/unit/openai-responses-subagent-strip-2446.test.ts. (thanks @like3213934360-lab)

  • fix(translator): GLM 5.2 (and other OpenAI-compatible upstreams that stream a tool call's id and function.name in separate SSE chunks) no longer produce an empty tool name / No such tool available: error through the Claude /messages path — the openai-to-claude streaming translator emitted content_block_start immediately on the id-only chunk with an empty name, and the Claude SSE protocol cannot patch a block after it is emitted, so the later name-only chunk was silently dropped. It now defers content_block_start until the tool name arrives (falling back to starting the block when arguments arrive first), so the emitted tool_use always carries the real name. Regression guard: tests/unit/openai-to-claude-glm-split-tool-name-2077.test.ts. (thanks @itiwant)

  • fix(resilience): OmniRoute didn't respect an exhausted Ollama Cloud (or any other apikey-category provider) quota — it retried the account seconds later instead of waiting out the real reset window (#6638) — shouldPreserveQuotaSignalsFor429()/checkFallbackError() (open-sse/services/accountFallback.ts) only applied body-text quota classification (daily/monthly/weekly quota-exhausted detection) to OAuth-category providers; apikey-category 429s (Ollama Cloud, OpenAI, etc.) always fell through to the generic short rate-limit cooldown regardless of what the error body said, and parseRetryFromErrorText() also had no support for day-granularity reset hints ("Your quota will reset in 3 days.") — only Xh/Ym/Zs combos. An explicit quota-exhausted signal in the body (looksLikeQuotaExhausted()) now overrides the apikey-category default via the new shouldPreserveQuotaSignals() (open-sse/services/quotaResetParsing.ts), and parseDayGranularityResetMs() parses whole-day reset countdowns so the real multi-day window is honored instead of a few seconds of backoff. Regression guard: tests/unit/issue-6638-ollama-quota.test.ts + 2 aligned tests/unit/account-fallback-service.test.ts cases that previously asserted the buggy rate_limit_exceeded/undefined-dailyQuotaExhausted behavior for apikey-provider quota text.

  • fix(resilience): a combo step "pinned" to one fingerprint account (mimocode/mcode/opencode multi-account providers) never actually resolved to that account, so it couldn't fail over when the pinned account was depleted (#6696, relates #6612) — the combo builder UI encodes an account pin as a composite connectionId (${rowId}|fp|${fingerprint}, src/lib/combos/builderOptions.ts), but expandTargetsByFingerprints() (open-sse/services/combo/fingerprintExpansion.ts) looked that composite string up directly in connectionById (keyed by real DB row ids), got undefined, and passed the target through unchanged, still carrying the bogus composite id — so downstream credential resolution could never match it either. expandTargetsByFingerprints() now splits the |fp| composite id back into the real connection row id + the pinned fingerprint (new splitFingerprintPin() helper) before any lookup, resolving the target to the real connectionId (with the pinned fingerprint carried on the new pinnedFingerprint field) instead of the inert composite string. Regression guard: tests/unit/combo-fingerprint-pin-6696.test.ts.

  • fix(api): Responses passthrough emitted event-only SSE frames (no data: line) for every dropped commentary event, breaking the OpenAI Python SDK's sse.json() parser (#6561), follow-up to #6199/#6232 — the commentary-drop continue; branches in open-sse/utils/stream.ts skipped the data: line for a dropped commentary event but never cleared the already-buffered event: line for that same frame, so the next blank line flushed the stale event: line alone. Both drop sites now call clearPendingPassthroughEvent() before continue, discarding the buffered prefix along with the dropped payload; the commentary-drop decision itself was extracted into a new open-sse/utils/responsesCommentaryDrop.ts so the fix does not grow the frozen stream.ts. Regression guard: tests/unit/responses-commentary-event-frame-6561.test.ts (realistic event:\ndata:\n\n frames — the existing #6199 test only used bare data: lines and never exercised this path).

  • fix(compression): /api/compression/preview's top-level originalTokens/compressedTokens diverged from engineBreakdown[0]'s counts for the same single-engine run (tiktoken outer counts vs the JSON.stringify(...).length/4 estimate per engine), worst on small inputs. A new reconcileSingleEngineTokens() overwrites the single-engine breakdown entry with the outer, more accurate figures; multi-step pipeline breakdowns are left untouched (#6488). Regression guard: tests/unit/compression/preview-outer-engine-token-reconcile-6488.test.ts.

  • fix(resilience): account selection could pick an account already out of quota upstream on every credentialed route except chat/codex (#6686) — getProviderCredentials() (src/sse/services/auth.ts) only skips a connection when a local cache already flags it exhausted (isQuotaExhaustedForRequest/src/domain/quotaCache.ts); it never itself calls the registered upstream QuotaFetcher. Only getProviderCredentialsWithQuotaPreflight() performs that live upstream check, and it was wired into exactly 2 call sites (src/sse/handlers/chat.ts, src/app/api/internal/codex-responses-ws/route.ts) — every other credentialed route (rerank, images/generations, images/edits, audio/transcriptions|speech|translations, videos/generations, music/generations, ocr, providers/[provider]/embeddings, providers/[provider]/images/generations, web/fetch, moderations, search) called the plain, cache-only selector, so an account whose cache entry was never populated (e.g. its first request landed on one of these routes) could be selected even at 0% quota remaining. Those 14 call sites now go through getProviderCredentialsWithQuotaPreflight() instead, matching chat/codex coverage. Regression guard: tests/unit/issue-6686-quota-preflight-coverage.test.ts (static check that none of the routes call the plain selector anymore + a behavioral check that the preflight-aware selector blocks a 100%-used account).

View originalPermalink
How v3.8.47 went

v3.8.46

Added 5
  • Hide paid-only models from auto/* routing when hidePaidModels setting is enabled
  • Add provider-family auto combos (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) that materialize virtual combos spanning installed backends for each model family
  • Implement native proxy-pool round-robin and egress IP rotation with multiple proxies per scope and rotation strategies (round-robin, random, sticky-per-N-min)
  • Add end-to-end tool and function calling support on the native Gemini /v1beta endpoint in both request and response directions
  • Add enterprise/work tier support for copilot-m365-web provider via M365_ENTERPRISE_OVERRIDES preset and agent field configuration
✨ New Features
  • feat(sse): hide paid-only models from auto/* routing when hidePaidModels is on (#6512) — follow-up to #6328/#6495. PR #6495 hid paid-only models from the GET /v1/models listing, but auto/* combos (auto/best-coding, auto/glm, …) could still pick a paid-only backend into their candidate pool → a 402/403 at request time. createVirtualAutoCombo now filters the candidate pool through the new pure open-sse/services/autoCombo/paidModelFilter.ts (filterPaidOnlyCandidates), applying the same free-model predicate #6495 uses in catalog.ts (providerHasFreeModels(provider) && isFreeModel(provider, {id})) whenever settings.hidePaidModels === true. Applied before the category/tier/family narrowing, so it covers every auto/* combo; an all-paid pool degrades to the existing graceful empty-pool path. Opt-in — default OFF leaves the pool unchanged (identity). Regression guard: tests/unit/autoCombo/paid-model-filter-6512.test.ts (4, incl. the default-off identity guard).
  • feat(sse): provider-family auto combosauto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini (#6453) — new routable ids that materialize an on-demand virtual combo spanning whatever installed backends currently expose that model family, degrading gracefully as backends rotate. A new pure open-sse/services/autoCombo/modelFamily.ts (detectModelFamily) classifies by model-id prefix for six families; zai is instead resolved by provider id (z.ai's hosted API serves the same glm-* model ids as every other GLM backend, so auto/zai means "route to my z.ai backend specifically" vs auto/glm's "any connected GLM backend"). Reuses the existing createVirtualAutoCombo on-demand materialization path (no DB writes) and the /v1/models catalog advertising loop. Regression guard: tests/unit/autoCombo/provider-family-combos.test.ts (11).
  • feat(proxy): native proxy-pool round-robin / egress IP rotation (#6365) — a scope (global / provider / account) can now hold multiple proxies as a pool with a rotation strategy, so outbound requests cycle their egress IP instead of pinning one proxy per scope. Migration 117_proxy_pool_rotation.sql lifts the UNIQUE(scope, scope_id) constraint (rebuild via the canonical rename/copy/drop; existing single assignments become 1-element pools) and adds a proxy_scope_rotation companion table holding the per-scope strategy + a persisted monotonic round-robin cursor. Strategies: round-robin (default, monotonic cursor — never Math.random), random, and sticky-per-N-min. Resolution (resolveProxyForScopeFromRegistry / resolveProxyForConnectionFromRegistry) now fetches the alive, position-ordered candidate set (unchanged PROXY_ALIVE_PREDICATE) and applies the strategy; an empty / all-dead pool still returns null — the #6246 fail-closed guard is untouched (never falls through to direct egress). Backend + DB only; dashboard pool-builder UI is a follow-up. Regression guard: tests/unit/proxy-pool-rotation-6365.test.ts (8, incl. fail-closed + backward-compat).
  • feat(providers): end-to-end tool/function calling on the native Gemini /v1beta endpoint (#6222) — both directions of the Gemini↔OpenAI conversion now preserve tool data (previously silently dropped). Request side: convertGeminiToInternal (extracted to its own testable module) maps tools[].functionDeclarations → OpenAI tools, prior functionCall parts → assistant tool_calls, and functionResponse parts → tool-role messages. Response side: convertOpenAIResponseToGemini emits parts[].functionCall {name,args} from message.tool_calls, and the streaming openAIChunkToGeminiChunk accumulates fragmented tool_calls deltas by index into complete functionCall parts. The non-Gemini client paths (Claude, OpenAI-Responses) already preserved tool calls — this closes the gap specific to the native Gemini surface. Regression guard: tests/unit/v1beta-gemini-tool-calling-6222.test.ts (6, incl. a streaming SSE round-trip).
  • feat(providers): copilot-m365-web enterprise / work tier support (#6334) — mirrors the EDU-tier pattern (#6210): M365ConnectionParams gains an agent field, a new opt-in M365_ENTERPRISE_OVERRIDES preset (agent=work, scenario=officeweb, licenseType=Premium) applies via providerSpecificData.tier="enterprise" (alias "work"), and agent is also overridable directly via providerSpecificData.agent. buildWsUrl was hardcoding agent="web" (the one enterprise-distinguishing param with no override path), so a Premium work account handshook then returned an empty stream. The individual and EDU paths are untouched. Kilo's dup flag vs #6210 (EDU tier) was a false positive — different tier. Regression guard: tests/unit/copilot-m365-enterprise-6334.test.ts (7). End-to-end confirmation on a real Premium work account is a live-VPS validation follow-up (Hard Rule #18). (thanks @Forcerecon)
  • feat(api): standardized, provider-agnostic effort + thinking request params (#6241) — a thin standardization layer over the existing mature per-provider reasoning plumbing (no provider mapper touched). providerChatCompletionSchema gains a canonical effort (reusing the shared none/low/medium/high/xhigh vocabulary — the UI tiers extra/max collapse onto xhigh) and a boolean thinking. A pure normalizeReasoningRequest (wired once in src/sse/handlers/chat.ts, before any reasoning field is read) folds them onto the fields the translators already consume (reasoning_effort / reasoning.effort / thinking), so they fan out to Anthropic / Gemini / xAI / Responses — an explicit client reasoning_effort / object-shaped thinking always wins (backward-compatible). /models additively exposes supportsThinking + effort_tiers so the frontend can render the toggles (UI component is a follow-up). Regression guard: tests/unit/effort-thinking-standardization-6241.test.ts (12). (thanks @Iammilansoni, @shabeer)
  • feat(combo): new pipeline (sequential) combo strategy (#6297) — the 18th routing strategy runs targets in order, threading each step's output into the next step's input, with an optional per-step prompt (system instruction); only the final step's response is returned. Distinct from fusion (parallel fan-out + judge). Implemented as a self-contained open-sse/services/pipeline.ts (sibling to fusion.ts), dispatched from combo.ts; the step list reuses combo.models order and reads an optional prompt off each target (backward-compatible — ignored by every other strategy). Intermediate steps run non-streaming with tools stripped (complete prose to thread forward); the final step keeps the client's stream flag + tools. A failing/empty/unparseable intermediate step fails the whole pipeline explicitly via a sanitized error (never silently swallowed). Kilo's dup flag vs #563 was a false positive (that's model→chain selection; this is a sequential chain). Regression guard: tests/unit/combo-pipeline-strategy.test.ts (5). (thanks @ofekbetzalel)
  • feat(ci): check:test-masking now flags inline-reimplemented prod conditions (#6348) — a new report-only subcheck (v2, 6A.10 family) catches the wrong-shape contract test: a test that recomputes the condition under test inline instead of importing/exercising the real function (the #6216 class, where === 500>= 500 stayed green because the test re-implemented the branch). For each added/modified test file it warns when the file textually duplicates a ≥3-token conditional from a production file touched in the same PR and does not import the symbol/module owning it, via a pure, fixture-tested findReimplementedConditions() with an allowlist mirroring assertReductionAllowlist. Report-only for now (does not fail the gate) — to be promoted to blocking after a triage cycle. Regression guard: tests/unit/check-test-masking.test.ts (45).
  • feat(sse): per-connection routing override (native vs CLIProxyAPI) (#6339) — the previously-dead isCliproxyapiDeepModeEnabled helper is now wired into resolveExecutorWithProxy: a single connection can opt itself into the CLIProxyAPI passthrough executor via providerSpecificData.cliproxyapiMode="claude-native", with precedence connection override > provider upstream_proxy_config mode > default. resolveExecutorWithProxy now receives the resolved connection's providerSpecificData (threaded from chatCore.ts), so one connection can deep-route while the provider's other connections stay native — no DB schema change (the toggle rides in providerSpecificData). Also resolves the same-provider mixing ask in #6340. Regression guard: tests/unit/chatcore-executor-proxy.test.ts (9). (thanks @RaviTharuma)
  • feat(dashboard): "Add session cookie" modal now shows a prominent "Open ‹host› →" link to the provider's own site (#6268) — every -web cookie-session provider (chatgpt-web, claude-web, gemini-web, kimi-web, lmarena, qwen-web, m365-copilot-web, …) renders a one-click external link (opening the provider's login/home page in a new tab) so operators no longer tab away to retype the URL mid-setup. The host resolves from a pure, unit-tested resolveWebProviderHost() (prefers WEB_COOKIE_PROVIDERS[id].website, falls back to the registry baseUrl origin); non-web providers render exactly as before. Kilo's dup flag vs #6265 (modal-too-small-on-1080p) was a false positive — distinct concern. Regression guard: tests/unit/resolve-web-provider-host.test.ts (5). (thanks @chirag127)
  • feat(providers): add DigitalOcean AI (serverless inference) as an OpenAI-compatible API-key provider (#6373) — base https://inference.do-ai.run/v1, wired through the shared OpenAI-compatible registry with full model passthrough (open-sse/config/providers/registry/digitalocean/, src/shared/constants/providers/apikey/inference-hosts.ts). Regression guard: tests/unit/digitalocean-provider.test.ts. (thanks @newnol)
  • feat(providers): add Huancheng Public API (hcnsec) as an OpenAI-compatible regional provider (#6410) — Xinjiang Huancheng Cybersecurity's public LLM platform (base https://api.hcnsec.cn/v1, free credits via daily check-ins), wired through the shared OpenAI-compatible registry with full model passthrough (open-sse/config/providers/registry/hcnsec/, src/shared/constants/providers/apikey/regional.ts). Regression guard: tests/unit/hcnsec-provider.test.ts. (thanks @UnrealAryan)
  • feat(dashboard): the web-session credential guide now shows an "Open {host}" link (#6316) to the provider's sign-in site (derived from the provider website via getProviderWebsiteHost), so you can jump straight to the page where the cookie/session must be captured. Regression guard: tests/unit/web-session-provider-link-6316.test.ts. (thanks @jordansilly77-stack)
  • feat(cerebras): add the Gemma 4 31B model (gemma-4-31b) to the Cerebras registry + pricing table (#6331). Regression guard extends tests/unit/t28-model-catalog-updates.test.ts. (thanks @backryun)
  • feat(providers): add Yuanbao (web) as a cookie-session provider (#6196) — yuanbao-web (Tencent Yuanbao, yuanbao.tencent.com) with cookie-only auth (hy_user/hy_token + public agent id), SSE→OpenAI translation incl. reasoning_content, exposing DeepSeek V3/R1 + Hunyuan / Hunyuan-T1. Regression guard: tests/unit/providers-yuanbao-web.test.ts. together-web was deferred (no verifiable web-session endpoint — needs a captured request) and huggingchat-web dropped (the existing huggingchat already is a web-cookie provider). (thanks @chirag127)
  • feat(providers): route the built-in agentrouter through the dynamic Claude-Code wire image (#6056) — a small static allow-set (CC_WIRE_IMAGE_BUILTINS in open-sse/services/ccWireImageBuiltins.ts), consulted by isClaudeCodeCompatible / isClaudeCodeCompatibleProvider / applyFingerprint, makes agentrouter adopt the CC wire-image headers + fingerprint while guarding the CC baseUrl/auth branches so it keeps its own registry baseUrl and x-api-key auth. Regression guard: tests/unit/agentrouter-cc-wire-image.test.ts (asserts the wire image is applied AND agentrouter's baseUrl/auth are preserved). Live WAF-acceptance against agentrouter.org is a VPS validation follow-up (Hard Rule #18).
  • feat(providers): bulk-add API keys for Cloudflare Workers AI (#6174) — cloudflare-ai is removed from the bulk-add exclusion list and the bulk parser gains a 3-field name|accountId|apiKey mode; the bulk route now builds a per-entry providerSpecificData so each key carries its own accountId (fixing the previous shared-object reuse), and both the create + key-validation paths receive it. Regression guard: tests/unit/bulk-api-key-parser-cloudflare.test.ts. (thanks @muflifadla38)
  • feat(dashboard): routing/settings UX clarity (#6147) — (1) weighted combos show the effective routing share % next to each weight when weights don't sum to 100 (WeightTotalBar.tsx); (2) the status widget's user-facing "Cloud Sync" label is renamed to "Remote Settings Sync" (CloudSyncStatus.tsx; internal ids/state untouched); (3) built-in providers gain an opt-in advanced base-URL override (isBaseUrlOverrideEligibleProvider, hidden behind an "Advanced" toggle, reusing the existing providerSpecificData.baseUrl persistence — not globally widened). Regression guard: tests/unit/routing-settings-ux-6147.test.ts.
  • feat(combo): add an option to disable session stickiness, per-combo or globally — round-robin / random combos can rotate to a different connection on every request instead of pinning a whole conversation to one connection by its first-message hash. Resolution precedence per-combo config.disableSessionStickiness → global settings.disableSessionStickiness → default false (preserves the #3825 prompt-cache/504 fix); gates both stickiness call sites in open-sse/services/combo.ts. Exposed as a global toggle (Combo Defaults) and a per-combo Inherit/on/off control. (#6168) Regression guard: tests/unit/combo-disable-session-stickiness.test.ts. (thanks @RCrushMe)
  • feat(docker): add the OMNIROUTE_NO_SUDO env flag for root-less / user-namespaced deployments — the MITM cert-trust command path (resolveSudoSpawn in src/mitm/systemCommands.ts) now strips the leading sudo when the flag is truthy, in addition to the existing root / sudo-missing cases, so the Proxy Agent runs without sudo (the operator trusts the CA manually, e.g. via NODE_EXTRA_CA_CERTS). Argv-array spawn preserved — no shell interpolation (Hard Rule #13). (#6122) Regression guard: tests/unit/mitm-systemCommands-no-sudo.test.ts. (thanks @powellnorma)
  • feat(providers): add Requesty as an OpenAI-compatible gateway provider (BYOK, base https://router.requesty.ai/v1, ~200 free requests/day) — wired through the shared OpenAI-compatible registry with full model passthrough (open-sse/config/providers/registry/requesty/, src/shared/constants/providers/apikey/gateways.ts). (#6120) Regression guard: tests/unit/requesty-provider.test.ts. (thanks @chirag127)
  • feat(dashboard): add configured-only / available-only filters to the Free Provider Rankings page (#6150) — hide providers you haven't configured, or whose connections are all rate-limited / out of quota, via server-side query params (?configuredOnly / ?availableOnly on GET /api/free-provider-rankings) backed by a testable lib helper reusing the in-process connection state (no Redis). Both filters default off, so the default view is unchanged; this supersedes the earlier client-side "Configured Only" toggle (#6245) with an available-only dimension and unit-tested logic. Regression guard: tests/unit/freeProviderRankings-filters.test.ts.
🔧 Bug Fixes
  • fix(dashboard): adding a second API key connection for the same provider no longer silently overwrites the first — the Add-API-key modal now derives a unique default connection name (main, then main-2, main-3, …) so the backend name-based upsert can't collide (#6499 — thanks @dilneiss).

  • fix(compression): the session-dedup engine now also deduplicates a large multi-line block repeated within a single message (intra-message dedup), not just across turns; the compression-preview API surfaces a fallbackReason, and the fusion panel reports how many models were rate-limited vs failed on total-panel failure (#6501 — thanks @chirag127).

  • fix(compression): a stacked-pipeline step naming an unregistered engine now surfaces a validationErrors entry instead of silently no-op'ing, so misconfigured pipelines are visible in the preview API (#6506 — thanks @chirag127).

  • feat(usage): add a Codex reset-credit redemption flow to the Provider Limits UI (#6361) — a useCodexResetCreditRedemption hook + /api/usage/codex-reset-credit route + codexResetCredits lib let you redeem banked Codex reset credits from the quota card. Regression guard: tests/unit/codex-reset-credits.test.ts. (thanks @JxnLexn)

  • feat(glm): add team-plan quota settings for glm-cn connections (#6351) — a dedicated GlmTeamQuotaFields form section (team quota id / limits) threaded through the Add/Edit connection modals, persisted via providerSpecificData, with the GLM usage service reading the team quota. Regression guards: tests/unit/glm-team-quota.test.ts, provider-specific-data-schema.test.ts. (thanks @hao3039032)

  • feat(providers): add TinyFish web-fetch/search support (#6349) — a tinyfish-fetch executor + /v1/web/fetch route + MCP web-fetch tool, registered as a specialty-media provider with request-validation and a search-provider catalog entry. Regression guards: tests/unit/executor-tinyfish-fetch.test.ts, web-fetch-handler.test.ts, mcp-web-fetch-tool.test.ts, provider-validation-tinyfish.test.ts. (thanks @dtybnrj)

  • fix(cli): omniroute launch-codex now spawns codex.cmd through a shell on Windows (the npm .cmd shim is unresolvable by bare spawn → ENOENT), mirroring the qodercli Windows fix (#6263) (#6312). Regression guard: tests/unit/launch-codex-windows-spawn-6312.test.ts. (thanks @swingtempo)

  • fix(codex): isolate the Spark quota from the shared Codex quota and stabilize the quota UI ordering / hydration so per-scope limits render consistently (#6336). Regression guards: tests/unit/codex-quota-selection-hydration.test.ts, provider-limits-ui.test.ts + 3 more. (thanks @xz-dev)

  • feat(api): add a hidePaidModels setting that filters paid-only models out of the /v1/models catalog. Regression guard: tests/unit/models-catalog-hide-paid.test.ts. (thanks @chirag127)

  • fix(api-manager): the fallback model picker now preserves combos instead of dropping them when a primary model is unavailable (#6443). Regression guard: tests/unit/api-manager-page-static.test.ts. (thanks @jmengit)

  • fix(providers): recoverable Antigravity / Cloud-Code (Gemini Code Assist) 403 responses (#6452) are now classified as a retryable project-config error instead of a terminal account ban, so a fixable project/API-disabled 403 no longer forces a ~1-year cooldown / full OAuth reconnect. Regression guard: tests/unit/errorclassifier-antigravity-403.test.ts. (thanks @developerjillur)

  • fix(mitm): sanitizeHeaders now redacts Set-Cookie response headers so upstream session cookies never leak into logs / diagnostics (#6451). Regression guard: tests/unit/mitm-sanitize-headers.test.ts. (thanks @developerjillur)

  • fix(api): /api/compression/preview now accepts mode: "caveman" and correctly handles stacked / zero-compression previews (#6425). Regression guard: tests/unit/api/compression-preview-caveman-and-stacked-6425.test.ts. (thanks @chirag127)

  • feat(providers): add Zed hosted LLM aggregator as a native-app provider (#6118) — OAuth sign-in via the Zed hosted flow, registered through the shared provider registry + executor. Regression guards: tests/unit/zed-oauth-provider.test.ts, zed-import-utils.test.ts, zed-docker-detect.test.ts, mitm-handler-zed.test.ts. VPS-validated via live operator login (Hard Rule #18).

  • fix(oauth): the Kiro SSO-cache auto-import now preserves the IDC region — cross-region Amazon Q / Kiro profiles imported from the SSO cache are no longer collapsed to the default region (#6113). Regression guard: tests/unit/kiro-auto-import-idc-2059.test.ts. VPS-validated via live operator login (Hard Rule #18).

  • fix(dashboard): passthrough model aliases no longer collide when two namespaced model ids share a last segment (port from 9router#1850, #6431). enx/gpt-5.5 and enx/codebuddy/gpt-5.5 both auto-generated the alias gpt-5.5, so the second model could never be added (the UI just alerted "alias already exists"). Aliases are now disambiguated deterministically — bare last segment when free, then parent-qualified (codebuddy-gpt-5.5), then a numeric suffix — while re-adding the exact same model id is still blocked. Regression guard: tests/unit/passthrough-alias-1850.test.ts. (thanks @arpicato)

  • fix(translator): preserve a Gemini functionResponse co-located with other parts (another functionCall, or trailing text) in the same content when translating Gemini → OpenAI (#6376). convertGeminiContent() early-returned the tool message on the first functionResponse part, dropping any co-located parts; such contents are now pre-split (one tool message per functionResponse, emitted first, plus one message for the remaining parts). Regression guard: tests/unit/gemini-to-openai-function-response.test.ts. (thanks @warelik)

  • fix(headroom): detect a python interpreter managed by mise / pyenv / asdf / conda (port from 9router#2353, #6382). Headroom's python probe (src/lib/headroom/detect.ts) searched a hardcoded PATH, but version managers expose their interpreters via shim dirs that only join PATH through interactive-shell activation — which the non-interactive server never runs, so a managed python (≥3.10) was invisible and Headroom reported it missing. The search path now prepends the well-known shim/bin dirs (~/.local/share/mise/shims, ~/.pyenv/shims, ~/.asdf/shims, $CONDA_PREFIX/bin, ~/.local/bin, respecting MISE_DATA_DIR/PYENV_ROOT/ASDF_DATA_DIR when set), and a new HEADROOM_PYTHON env override lets operators point straight at their interpreter (mirroring HEADROOM_URL). Still shell-free (execFileSync). Regression guard: tests/unit/headroom-detect.test.ts (5). (thanks @loopyd)

  • fix(executors): strip the OpenAI-Codex/Claude-CLI client_metadata passthrough field for NVIDIA requests (port from 9router#1887, #6411). NVIDIA's OpenAI-compatible wrapper rejects it with 400 Unsupported parameter, the same class already handled for cerebras/mistral; nvidia (executor default) was missing from the strip allowlist so Codex/Claude-Code passthrough requests 400'd. Regression guard: tests/unit/executor-default-strip-client-metadata.test.ts (+nvidia case). (thanks @phidinhmanh)

  • fix(translator): strip the Claude-style thinking field for NVIDIA z-ai/glm-5.2 (port from 9router#2023, #6413). NVIDIA's OpenAI-compatible wrapper 400s on thinking (a Claude-format client routed here leaves a thinking:{type:"adaptive"}); the existing strip rule only dropped reasoning. Same class already handled for minimax-m2.7. Regression guard: tests/unit/nvidia-minimax-thinking-strip.test.ts (+glm-5.2 case). (thanks @phidinhmanh)

  • fix(translator): suppress the streamed </think> close marker for the Antigravity IDE client (port from 9router#1061, #6415). On thinking-only turns Antigravity rendered a bare </think> as the sole visible content, tripping its loop-detection and wasting requests. Antigravity's UA (vscode/<v> (Antigravity/<v>)) is added to the marker-suppress allowlist (alongside OpenCode); Claude Code / Cursor still get the marker, and x-omniroute-thinking-marker: on force-restores it. Regression guard: tests/unit/think-close-marker-suppress-5245.test.ts. (thanks @abdofallah)

  • fix(executors): strip nested reasoning_content from messages for Mistral (port from 9router#1649, #6417). Mistral's API returns 422 extra_forbidden when an assistant message carries reasoning_content (replayed thinking from a prior turn, e.g. via the Codex /responses path); the generic top-level 400 field-downgrade retry never covered the nested per-message field. DefaultExecutor now strips it for provider mistral only, so DeepSeek (which requires replayed reasoning_content) is unaffected. Regression guard: tests/unit/mistral-strip-reasoning-content-1649.test.ts. (thanks @xxy9468615)

  • fix(executors): strip the client_metadata passthrough field on the OpenCode path (port from 9router#1442, #6418). OpenCode upstreams (e.g. kimi-k2.6 via opencode-go) reject it with 400 "Extra inputs are not permitted, field: 'client_metadata'"; the DefaultExecutor strip only covered cerebras/mistral and OpencodeExecutor extends BaseExecutor directly, so nothing removed it there. Regression guard: tests/unit/opencode-strip-client-metadata-1442.test.ts. (thanks @yanpaing007)

  • fix(executors): inject the reasoning_content echo for the native Moonshot Kimi provider (port from 9router#1480, #6419). Kimi (executor default) is a thinking-mode upstream that 400s with "reasoning_content must be passed back" when a prior assistant turn lacks it; the placeholder injection was only wired into the OpenCode meta-provider, so direct multi-turn Kimi conversations failed. Scoped to kimi (gateway-served models matching the thinking-model name pattern are unaffected). Regression guard: tests/unit/kimi-native-reasoning-injected-1480.test.ts. (thanks @2220258345)

  • fix(executors): recover from a strict gateway's context_management: Extra inputs are not permitted 400 (port from 9router#1468, #6420). Claude Code always sends a top-level context_management field; strict anthropic-compatible gateways reject it. The dedicated context-editing 400-fallback only fired when OmniRoute's own contextEditing feature was enabled (default off), so a client-sent field passed through untouched and 400'd. context_management is now in the generic reactive field-strip list, so it's stripped-and-retried once regardless of the feature flag (with correct request re-signing for claude-compatible relays). Regression guard: tests/unit/provider-field-strips.test.ts. (thanks @ohahe52-dot)

  • fix(network): enable RFC 8305 Happy Eyeballs (autoSelectFamily) on the direct-egress undici dispatcher (port from 9router#1237, #6423). When DNS returns both IPv6 (AAAA) and IPv4 (A) and the IPv6 route is broken (e.g. a NAT64 64:ff9b:: prefix without routing), undici tried IPv6 first and hung until ETIMEDOUT (then a 502 + account lockout), even though curl reached the same host. The direct dispatcher now races both families and uses whichever connects first. Proxy paths pin family via proxyTls and are unaffected. Regression guard: tests/unit/direct-dispatcher-pipelining-4580.test.ts. (thanks @adentdk)

  • fix(combo): round-robin now advances the rotation pointer past the model that actually served, not the eagerly-scheduled one (port from 9router#948, #6428). With stickyLimit: 1 (true round-robin), when the scheduled model failed and a different model served via fallback, the counter had already advanced +1 from the scheduled index — so the next request reused the fallback-served model, degrading round-robin into hot-spotting on whichever model was healthy. The pointer now advances to the served index + 1 (mirroring the sticky-limit>1 path). Session-stickiness (#3825) and distribution are preserved. Regression guard: tests/unit/combo-rr-fallback-advance-948.test.ts. (thanks @binsarjr)

  • fix(sse): a non-string model field is now rejected with a 400 before the resolver, instead of crashing downstream .toLowerCase()/.split() calls into an empty-body 500 that escapes the error sanitizer (#6407). Regression guard: tests/unit/chat-non-string-model-6407.test.ts. (thanks @chirag127)

  • fix(api): unknown /api/* routes now return a JSON 404 (instead of the dashboard HTML shell) and scalar chat params (model/temperature/etc.) are validated before the provider lookup so malformed requests fail fast with a clear 400 (#6424, #6412). Regression guards: tests/unit/api/api-catchall-json-404.test.ts, tests/unit/chat-early-schema-validation-6412.test.ts. (thanks @chirag127)

  • fix(api): /v1/chat/completions now rejects a non-JSON Content-Type with a 400 before parsing the body (#6414). Regression guard: tests/unit/v1-chat-completions-content-type-6414.test.ts. (thanks @chirag127)

  • fix(api): the X-OmniRoute-Compression response header is now echoed on /v1/chat/completions and /v1/completions (#6422). Regression guard: tests/unit/compression-header-echo-6422.test.ts. (thanks @chirag127)

  • fix(api): concurrent GET /v1/models requests are coalesced into a single catalog build (#6408). Regression guard: tests/unit/v1-models-concurrent-6408.test.ts. (thanks @chirag127)

  • fix(api): /v1/completions now echoes the requested body.model in its JSON + streamed responses (#6429). Regression guard: tests/unit/completions-body-model-echo.test.ts. (thanks @chirag127)

  • fix(api): env-var master keys now see the full /v1/models catalog (#6406). Regression guard: tests/unit/models-catalog-envkey-6406.test.ts. (thanks @chirag127)

  • fix(api): non-streaming /v1/completions responses now echo body.model aligned with the X-OmniRoute-Model header (#6426). Regression guard: tests/unit/v1-completions-model-header-match-6426.test.ts. (thanks @chirag127)

  • fix(api): unknown /v1/* routes now return a JSON 404 not_found instead of the Next.js dashboard HTML shell (#6405). Regression guard: tests/unit/api/v1-catchall-json-404.test.ts. (thanks @chirag127)

  • fix(api): the per-connection provider models route now degrades to the shipped catalog when a provider's /models endpoint answers with a redirect (#6267) — a qwen-web import failed with a raw Redirect blocked … (307) 503. safeOutboundFetch throws REDIRECT_BLOCKED on the 307, getSafeOutboundFetchErrorStatus maps it to 503, and buildDiscoveryErrorFallbackResponse treated every 503 as a hard error — so the non-empty getModelsByProviderId("qwen-web") catalog was never surfaced. A models-endpoint redirect is not a fixable-config error (unlike URL_GUARD_BLOCKED/INVALID_URL, which stay hard errors), so it now falls back to the local/cached catalog before the 503 short-circuit. General fix — covers any config-driven provider that 307s. Regression guard: tests/unit/provider-models-qwen-web-redirect-6267.test.ts. (thanks @chirag127)

  • fix(api): the per-connection provider models route (MCP list_models_catalog + the dashboard import view) now merges USER-ADDED custom models into its response (#6247) — custom models live in the key_value namespace customModels, which the live REST /api/v1/models already merges, but src/app/api/providers/[id]/models/route.ts never read getCustomModels, so custom models were dropped on both the discovery-success and local_catalog paths. They are now folded into the returned model list (deduped by id, stamped owned_by: provider), fixing MCP + the dashboard import view in one place. Regression guard: tests/unit/provider-models-custom-merge-6247.test.ts. (thanks @RCrushMe)

  • fix(providers): GitLab Duo tool-calling follow-up turns no longer fail upstream with 422 {"detail":"Validation error"} (tokens 0/0, rejected pre-inference). The #6234 tool-result-feedback fix serialized the entire multi-turn conversation into GitLab's single-file code_suggestions (small_file) generation endpoint — folded history that turn-N sent as an oversized current_file.content_above_cursor and duplicated verbatim into user_instruction, tripping the AI-Gateway's small_file validation guard. The executor now bounds that prompt: it keeps system + latest user message + the most-recent tool round (dropping older turns), caps oversized tool results, and stops duplicating the full prompt into user_instruction (which now carries only the short latest user message) — while still feeding the most-recent tool result back so the agent continues (open-sse/executors/gitlab.ts, #6220). The unit test covers the bounding logic; the upstream 422→200 clearing is VPS-only (Hard Rule #18). Regression guard: tests/unit/gitlab-tool-exchange-bounded-6220.test.ts.

  • fix(i18n): the provider-detail (/dashboard/providers/[id]) connection-status filter labels no longer render as __MISSING__:All / __MISSING__:Active / __MISSING__:Error / __MISSING__:Banned / __MISSING__:CreditsExhausted in non-English locales (notably pt-BR) (#6290). Root cause was not the namespace mismatch the issue guessed — the providers.filter* keys resolve correctly in en.json; the debt lived in the locale mirrors (src/i18n/messages/*.json), where these five keys carried the __MISSING__: sync sentinel in ~15 locales and were absent entirely in ~26 others, so next-intl found the key and echoed the sentinel verbatim. All 40 non-English/-Chinese mirrors now ship real translations for the five providers.filter* labels. Regression guard: tests/unit/i18n-provider-filter-keys-6290.test.ts. (thanks @diegosouzapw)

  • fix(providers): the copilot-m365-web streaming executor now emits debug-level WebSocket diagnostics (#6210) — the outbound WS URL (with the access_token redacted via redactWsUrl()), handshake success/failure, and each received SignalR frame's type/target. Previously the streaming path logged nothing, so an empty content:null response (the M365 Education / Starter tier symptom fixed in #6234) was undiagnosable even at APP_LOG_LEVEL=debug. The change is debug-level and side-effect-free — it does not alter streaming behavior or the frame parser, and the token never reaches the logs. Regression guard: tests/unit/copilot-m365-web-logging-6210.test.ts (thanks @qpeyba)

  • fix(resilience): a round-robin combo no longer returns 503 all upstream accounts are unavailable when a compatibility-rejected target is actually healthy (#6238). filterTargetsByRequestCompatibility drops request-incompatible targets (tool/vision/structured-output unsupported, or below the required context window) before any availability check runs, and its compatible.length === 0 safety net only fired when all targets were filtered — not when the kept targets later all turned out runtime-unavailable (circuit-open / cooldown / no credentials). So a combo could 503 while a compat-rejected-but-healthy provider sat unused. handleRoundRobinCombo now keeps the compat-rejected set and, when every compat-kept target was skipped without a single real attempt, probes those rejected targets as a last-resort fallback tier (via the new pure open-sse/services/combo/comboCompatFallback.ts) before crystallizing the 503. Regression guard: tests/unit/combo-roundrobin-compat-fallback-6238.test.ts. (thanks @ThongAccount)

  • fix(startup): best-effort self-heal for a corrupted Turbopack dev cache on Windows (#6289). On Windows, pnpm dev can fail at startup when Turbopack mmaps a persistent-cache SST file and the OS refuses the mapping (os error 1455 — "paging file too small"), which Turbopack surfaces as a misleading Module not found: Can't resolve '@/shared/utils/machine'. This is a known upstream Turbopack cache-corruption bug — not our code. The dev launcher (scripts/dev/run-next.mjs) now wraps nextApp.prepare() and, when it rejects with that signature (isTurbopackCacheCorruption in the new scripts/dev/turbopackCacheHeal.mjs), purges .build/next/**/cache/turbopack and retries once with a clear log. Caveat — best-effort only: the corruption often surfaces as a runtime overlay rather than a prepare() rejection, so this cannot always intercept it; the reliable remedy remains manually deleting the Turbopack cache dir. Regression guard: tests/unit/turbopack-cache-heal-6289.test.ts. (thanks @chirag127)

  • fix(providers): qodercli PAT auth no longer fails with spawn qodercli ENOENT on Windows (#6263) — spawnQoderCli spawned the bare qodercli name with shell:false and an unenriched env, so the npm .cmd wrapper under %APPDATA%\npm (a user-PATH directory) was never resolved. It now resolves the absolute .cmd/.exe path through the existing getCliRuntimeStatus("qoder") resolver in src/shared/services/cliRuntime.ts (memoized), spawns with shell when the target is a .cmd/.bat, and uses the cliRuntime-enriched env (PATH + PATHEXT + APPDATA); the ENOENT error now lists the searched paths plus the CLI_QODER_BIN override. End-to-end spawn on a real Windows host is host-only (Hard Rule #18); the path-resolution logic is unit-tested. Regression guard: tests/unit/qodercli-windows-resolve-6263.test.ts. (thanks @chirag127)

  • fix(sse): the reasoning-token buffer no longer inflates probe-sized max_tokens (#6274) — Claude Code's /model capability check sends max_tokens: 1, but for a thinking-capable model with a large output cap (e.g. glm-5.2) the #3587 headroom heuristic (max(current + 1000, ceil(current * 1.5))) rewrote it to 1001 and forwarded that upstream, wasting tokens on a request that was never a genuine reasoning budget. resolveReasoningBufferedMaxTokens() (open-sse/services/reasoningTokenBuffer.ts) now short-circuits and returns the caller's value verbatim when it is below the new REASONING_BUFFER_MIN_TRIGGER (256) threshold — a tiny explicit limit is a probe, not a reasoning request. Real budgets still receive the #3587 headroom unchanged, and the guard runs after the existing capability checks so unknown / non-reasoning models keep returning null. Regression guard: tests/unit/reasoning-token-buffer-6274.test.ts. (thanks @brightfiscalband)

  • fix(cli): omniroute reset-password now works as a real subcommand, and password resets over piped (non-TTY) stdin actually apply (#6261, #6258). Two coupled defects: (1) #6261bin/omniroute.mjs routed everything through Commander with only two pre-Commander bypasses (--mcp, reset-encrypted-columns), so omniroute reset-password was rejected as an unknown command; only the separate omniroute-reset-password bin worked, while the docs falsely advertised the subcommand (incl. a bogus "legacy alias still works"). A pre-Commander bypass mirroring reset-encrypted-columns now dynamically imports bin/reset-password.mjs (which self-executes) before Commander parses; the three doc lines were corrected. (2) #6258bin/reset-password.mjs issued two sequential rl.question prompts; under piped stdin the second read never settled at EOF, so main() never reached resetManagementPassword and the reset was a silent no-op (both prompts printed, no success, password unchanged). The CLI now detects non-TTY stdin and reads it once (first line = password, second line = confirm if present, else reused), adds a --password-stdin flag (entire stdin is the password, no confirmation), and exits 0 explicitly so the success line always flushes; interactive TTY behavior is unchanged. Regression guard: tests/unit/reset-password-cli-6261-6258.test.ts (3). (thanks @chirag127)

  • fix(db): the mass-migration safety abort now tells the operator how to bypass it and stops flooding the log (#6260) — after restoring a backup that wiped the migration tracking table, runMigrations() threw the abort on every downstream ensureDbInitialized(), re-logging the full banner 11+ times, and the message never mentioned the existing OMNIROUTE_MAX_PENDING_MIGRATIONS escape hatch. The abort text now appends a bypass hint (set OMNIROUTE_MAX_PENDING_MIGRATIONS=0 in server.env / DATA_DIR/.env), and a new MigrationSafetyAbortError is memoized so repeated calls in the same process throw the same instance and emit a single concise line instead of the full cascade. Regression guard: tests/unit/migration-safety-abort-6260.test.ts. (thanks @chirag127)

  • fix(auth): importing a distinct Codex/ChatGPT OAuth auth.json is no longer falsely rejected as "already exists" when it belongs to a different user in the same workspace (#6301). findExistingCodexConnection (in src/lib/oauth/utils/codexAuthImport.ts) deduped only on providerSpecificData.workspaceId === accountId, where accountId is the shared chatgpt_account_id/tokens.account_id — so two members of the same ChatGPT Team collapsed onto a single connection (409 duplicate_account). The id_token's https://api.openai.com/auth claim carries a per-user chatgpt_user_id alongside the workspace id (the device-flow path already persisted it as chatgptUserId, but the import path did not). Now parseAndValidateCodexAuth extracts userId (chatgpt_user_iduser_id → JWT sub) into ParsedCodexAuth, the create/update paths persist chatgptUserId in providerSpecificData (mirroring codex.ts), and dedup keys on workspaceId AND chatgptUserId — with a backward-compat fallback to legacy accountId-only matching when no stored connection for that workspace records a chatgptUserId, so genuinely-same accounts still dedup. Regression guard: tests/unit/codex-auth-import-userid-dedup-6301.test.ts (4). (thanks @anungma)

  • fix(providers): importing models for the venice-web provider no longer fails with a red "Provider venice-web does not support models listing" (#6269). venice-web is a web-cookie provider with an executor but no upstream /v1/models endpoint and no registry models, so the models route fell through to the tail 400. Mirroring the jules/linkup-search/ollama-search fix (#5569), it now ships a static local catalog entry in src/lib/providers/staticModels.ts — seeding the current Venice lineup (venice-uncensored, llama-3.3-70b, qwen3-235b, qwen3-4b, deepseek-r1-671b; Venice rotates its catalog, see docs.venice.ai/models/overview) — so the route returns 200 with source:"local_catalog", intentional:true. Regression guard: tests/unit/static-models-venice-web-6269.test.ts. (thanks @chirag127)

  • fix(api): the specialty model catalogs (/v1/embeddings, /v1/images, /v1/music, /v1/videos model lists) are now derived from the unified catalog filtered by a predicate (getSpecialtyModelsResponse) instead of ad-hoc per-route logic, so they consistently respect active-credential visibility and stay in sync with the main catalog (#6303). Regression guard: tests/unit/specialty-model-catalog-routes.test.ts. (thanks @makcimbx)

  • fix(api): the agent-bridge server route now resolves the MITM manager via a dynamic import("@/mitm/manager.runtime") so Turbopack does not statically pull the stub (or over-bundle the manager), and the agent-skills generator anchors its output base path with path.join(process.cwd(), …) so Turbopack's static analyzer stops tracing the whole project root (#6329, #6366). Regression guard: tests/unit/agent-bridge-server-route-dynamic-import.test.ts. (thanks @Iammilansoni)

  • fix(api): internal probes (combo-test, cloud-sync verify) now pick a management-scoped / allow-all API key instead of naively grabbing getApiKeys()[0] — a restricted self:usage first row made the probe fail with "Model X is not allowed for this API key" even when the combo path was healthy (pickApiKeyForInternalUse in src/lib/db/apiKeys.ts). The API-manager model editor also falls back to /api/models?all=true when /v1/models is catalog-protected (#6372). Regression guard: tests/unit/pick-internal-api-key-6372.test.ts. (thanks @jmengit)

  • fix(live-ws): the Live Dashboard WebSocket server now rejects on bind failure (e.g. EADDRINUSE when the API bridge already holds the port) instead of letting the error surface as an unhandled error event that crash-loops the process — the error listener is attached to wss (not server) and releases the EventBus subscription on a failed start (#6324). Regression guard: tests/unit/live-ws-eaddrinuse-6324.test.ts. (thanks @vinayakkulkarni)

  • fix(dashboard): the Home provider-topology widget now trusts the live provider-metrics snapshot — it uses topology.errorProvider and live activeRequests directly instead of re-deriving state from a stale lastErrorAt or applying a frontend timeout filter, so the topology reflects real-time provider health (#6322). Regression guard: tests/unit/home-provider-topology-live-state.test.ts. (thanks @xz-dev)

  • fix(sse): strip zero-width markers from streamed tool-call arguments — a follow-up to #5857. That PR removed injected zero-width joiners (U+200D) from streamed assistant text/reasoning but deliberately left tool-call argument JSON byte-exact. The request-side obfuscation (open-sse/services/claudeCodeObfuscation.ts) injects ZWJ into agent words — including the temp path inside the Bash tool description — and Claude models copy that verbatim into generated commands, which are delivered as tool-call arguments rather than assistant text. As a result the ZWJ survived and corrupted code blocks (e.g. a temp path rendered with an invisible joiner). Now open-sse/handlers/responseSanitizer.ts strips zero-width code points from tool-call argument strings at every emit site (OpenAI non-stream/stream chat tool_calls + legacy function_call, native Responses function_call items, the OpenAI→Responses conversion, and the native Responses streaming response.function_call_arguments.delta/.done events). Only zero-width code points are removed; JSON structure and all other bytes stay identical (no parse/restringify), so normal arguments remain byte-exact. Regression guard: 6 new cases in tests/unit/response-sanitizer.test.ts (suite 50/50).

  • fix(nodejs): the default app log path now resolves under DATA_DIR (~/.omniroute/logs/application/app.log) instead of process.cwd() (#6197) — the globally-installed CLI runs from an arbitrary working directory, so anchoring the default to cwd made file logging silently write to (or no-op under) an unrelated directory, contradicting the documented .env.example default. getAppLogFilePath() now computes the default lazily via the pure resolveDataDir() resolver (honours a per-process DATA_DIR, no directory-creation side effect); an explicit APP_LOG_FILE_PATH still wins. Regression guard: tests/unit/logenv-datadir-path-6197.test.ts (3).

  • fix(docker): AgentBridge/startMitm no longer aborts in containers/headless when the Antigravity-default DNS step can't write /etc/hosts (#6127), and the privileged command's stderr now reaches app.log instead of only a bare exit code hitting the toast (#6198). The default DNS step (addDNSEntry) was called unguarded while cert install and the two sibling DNS steps were each best-effort — in the runtime Docker image (USER node, no sudo, read-only /etc/hosts) it threw Command failed with code 1 out of startMitmInternal and killed the whole start, discarding the stderr. The three DNS steps are extracted into a best-effort provisionDnsEntries() where each failure is logged with the full err (stderr included, folded in by systemCommands.ts) and never aborts the start. Regression guard: tests/unit/mitm-dns-graceful-degrade-6127.test.ts (4).

  • fix(providers): copilot-m365-web now supports the M365 Education "Starter / OfficeWebIncludedCopilot" tier and no longer returns an empty content:null stream (#6210). Two gaps: (1) buildWsUrl() hardcoded the individual-consumer scenario (OfficeWebPaidConsumerCopilot, isEdu=false) — the EDU tier is now opt-in via providerSpecificData.tier="edu", emitting scenario=OfficeWebIncludedCopilot/isEdu=true (the individual path is unchanged); (2) the EDU/GPT-5.5 path streams deltas via arguments[0].writeAtCursor (incremental) instead of only messages[].text (accumulated snapshots), which the parser dropped — a new accumulateBotContent() folds both formats, with type:2 item.result.message as a last-resort fallback. Regression guard: tests/unit/copilot-m365-edu-writeatcursor-6210.test.ts (10). (thanks @qpeyba)

  • fix(providers): GitLab Duo executor now feeds tool results back into the prompt instead of looping (#6220) — buildPrompt() branched only on system/user and took userParts.at(-1), silently dropping the assistant{tool_calls} + tool{result} turns the client appended, so the reconstructed prompt was byte-identical to turn 1 and the model re-emitted the same <tool> call forever. When a tool exchange is present the full conversation is now serialized, folding each tool result back keyed by its tool_call_id; simple conversations keep the legacy shape. Complements the tool_call emission from #6051 (the kilo-duplicate label was a false positive — different, sequential defect). Regression guard: tests/unit/gitlab-tool-result-feedback-6220.test.ts (4).

  • fix(providers): opencode-go/opencode-zen can now synthesize the OpenCode CLI identity headers Cloudflare requires on VPS egress (#5997) — on a datacenter VPS, opencode.ai/zen/go/v1/chat/completions 403s (HTML challenge) requests lacking CLI identity, while the reporter's control curl proved that User-Agent: opencode-cli/1.0.0 + x-opencode-client: cli + x-opencode-project: default + fresh request/session UUIDs succeed. Opt-in via OPENCODE_SYNTHESIZE_CLI_HEADERS=true (values overridable via OPENCODE_GO_USER_AGENT/OPENCODE_USER_AGENT/OPENCODE_CLIENT/OPENCODE_PROJECT); it fills only headers the client did not already send. Kept off by default — the forward-only path is deliberate (fabricating a wrong value risks upstream rejection; a prior dedup regressed with opencode/local), so this replaces the fragile local header-injection shim without changing default behavior. Regression guard: tests/unit/opencode-cli-headers-synthesis-5997.test.ts (6). (thanks @aleksesipenko)

  • fix(resilience): sticky session affinity now evicts and fails over to another account when the pinned account is exhausted/unavailable (#6219)

  • fix(sse): Responses API passthrough now drops internal commentary-phase output before forwarding to clients (gated by RESPONSES_PASSTHROUGH_DROP_COMMENTARY, default on) (#6199)

  • fix(sse): tool-call function schemas with a root type: null are now coerced to type: "object" before dispatch (#6359) — clients like the Codex app emit parameters: { type: null, ... } for some tools, which OpenAI-compatible upstreams reject with 400 Invalid schema for function '...': schema must be a JSON Schema of 'type: "object"', got 'type: null', failing the whole request. toolSchemaSanitizer already stripped the null; it now re-adds the mandatory root "object" type (and empty properties/open additionalProperties when absent). Combinator roots (anyOf/oneOf/allOf) and explicit root types are left untouched. Regression guard: 5 new cases in tests/unit/tool-schema-sanitizer.test.mjs.

  • fix(docker): AgentBridge no longer fails to start on npm/Electron/VPS installs with "MITM manager stub reached at runtime" (#6344) — v3.8.45 flipped the production bundler default to Turbopack, but next.config.mjs aliased @/mitm/manager to its Docker-only degraded stub unconditionally. That was harmless while Docker (which sets the alias intentionally for #3390 graceful degradation) was the sole Turbopack consumer, but once every artifact built with Turbopack the stub shipped to all non-Docker users and startMitm threw on the first Agent-Bridge start. The alias is now opt-in via OMNIROUTE_MITM_STUB=1 (set only by the Dockerfile) through the shared scripts/build/mitm-stub-flag.mjs helper; default builds bundle the real manager. Regression guard: tests/unit/mitm-stub-alias-6344.test.mjs (4).

  • fix(proxy): stop the v3.8.44 proxy regression that leaked the real IP and disabled healthy proxies (#6246). Two coupled defects from the new health scheduler: (1) IP leak — when a proxy assigned to a connection was marked inactive, resolution fell through to a direct egress instead of blocking, exposing the operator's real IP; (2) over-deactivation — the sweep flipped a proxy to inactive on the first failed probe and counted our own 5s timeout / a probe-target 5xx as the proxy's fault, so healthy paid proxies vanished from egress selection ("my proxies are not being used anymore"). Fix: the sweep decision is extracted into a pure, network-free decideProxyHealthAction (src/lib/proxyHealth/decision.ts) — by default the health check now only counts/logs and never downgrades status (a proxy is downgraded/removed only with PROXY_AUTO_REMOVE=true, after PROXY_AUTO_REMOVE_AFTER consecutive conclusive failures); probes are classified tri-state so an inconclusive result (our timeout, or a 5xx from the probe target) never penalizes the proxy, and the probe timeout is raised 5s→15s. Separately, safeResolveProxy now fails closed via the existing policy: a connection whose assigned proxy is dead is blocked instead of leaking direct (hasBlockingProxyAssignment), honoring the explicit proxy off toggles and the PROXY_FAIL_OPEN=true opt-out. Existing proxies stuck inactive by the old behavior need a one-time manual re-activate (the operator owns proxy status). Regression guards: tests/unit/proxy-health-decide-action-6246.test.ts, tests/unit/proxy-assigned-unavailable-6246.test.ts.

  • fix(proxy): make "Test All" read-only and add bulk enable/disable (#6246). Complements the core fail-closed / scheduler fix (#6296) with the two remaining reporter asks. (1) The "Test All" button (POST /api/settings/proxies/auto-test) used to flip a proxy to inactive on a failed reachability probe; since the egress selector excludes inactive proxies, a flaky probe (an unreachable httpbin.org, a proxy that blocks HEAD, or a slow paid proxy) silently disabled every proxy that failed — "Test All" is now read-only by default (only the operator sets a proxy active/inactive; opt back into the legacy test-and-set with PROXY_HEALTH_AUTO_DEACTIVATE=true). (2) Adds a bulk enable/disable proxies endpoint + toolbar action (POST /api/settings/proxies/batch-activate) so an operator can re-activate proxies in one click. Regression guard: tests/unit/proxy-health-6246.test.ts. (thanks @tenshiak)

  • chatcore (tools): stop the default 128-tool cap from silently dropping opencode's task/MCP tools. opencode (used as an MCP/agent host) sends a large tool list; when it exceeds the speculative MAX_TOOLS_LIMIT (128) default, truncateToolList did a blind tools.slice(0, 128), dropping every tool past index 128 — including opencode's built-in task tool (subagent launch) and many MCP tools, so models routed through OmniRoute could no longer spawn subagents or reach part of their tools. The cap exists to avoid upstream 400s for providers with real hard limits (e.g. grok-cli 200), so it is kept for those: detection of the opencode client (isOpencodeClient — any x-opencode-* header, or opencode in the user-agent) now only bypasses the speculative 128 default, never a known provider ceiling. Precedence is explicit — a proactive/detected provider limit always truncates (even for opencode); otherwise opencode forwards its full tool list; otherwise the unchanged 128 default applies to every other client. Refactors getEffectiveToolLimit into getKnownToolLimit(provider) ?? DEFAULT_LIMIT (byte-identical for existing callers) and fixes a cosmetic debug-log that reported the truncated count instead of the original. Regression guard: tests/unit/tool-limit-detector.test.ts.

  • fix(mitm): the macOS MITM-cert install check now matches the system keychain again. security find-certificate -a -Z prints the SHA-1 as a colon-less hex string, but the installed-check compared it against getCertFingerprint()'s colon-separated form, so the substring match never hit — the cert was reported as not-installed and re-prompted for the sudo install on every run. Fingerprints are now normalized (colons stripped, upper-cased) on both sides via the extracted macCertOutputHasFingerprint helper. Regression guard: tests/unit/mitm-cert-mac-fingerprint.test.ts. (#6204, closes #6134 — thanks @rianonehub)

  • fix(api): /v1/messages/count_tokens now counts tool_use, tool_result and thinking content blocks (and array-form system prompts) in the local-estimation path, instead of only text. Real agentic conversations keep ~95% of their tokens inside tool results; the previous estimate returned near-zero for them, which silently broke Claude Code's auto-compaction (context grew past the window with no compaction until the upstream API rejected the request). The real provider-side count path is unchanged. Regression guard: tests/unit/messages-count-tokens-route.test.ts. (#6221 — thanks @luweiCN)

  • fix(antigravity): strip a trailing assistant prefill turn for Vertex Claude models to avoid upstream 400s (#6114). Regression guard: tests/unit/antigravity-claude-prefill-strip.test.ts. (thanks @anki1kr)

  • fix(security): the mutable cloud-agent routes (/api/cloud/credentials/update, /api/cloud/models/alias) now require management auth instead of being treated as public. They were classified as public API routes, so a request without management credentials could update stored cloud-agent credentials and model aliases. They are removed from the public-route set, classified as management routes in the authz pipeline, and gated by requireManagementAuth; cloud read/auth routes stay public. Regression guards: tests/unit/cloud-write-auth.test.ts, tests/unit/authz/classify.test.ts, tests/unit/public-api-routes.test.ts. (#6233 — thanks @vittoroliveira-dev)

  • refactor(dashboard): extract the onboarding-wizard "Open provider details" link target into a pure, unit-tested buildProviderDetailsHref(connection) helper. The wizard already routes by connection.id (the node UUID) rather than the provider category slug (#6144/#6145); this hardens that behavior behind a tested helper that guards a missing id/connection. Regression guard: tests/unit/provider-onboarding-href.test.ts. (#6166 — thanks @KooshaPari)

  • fix(security): the doubao synthetic device-id generator now derives its digits via an unbiased crypto-random draw (rejection sampling over crypto.randomBytes()) instead of a % 10 reduction, closing a CodeQL js/biased-cryptographic-random finding.

  • fix(agentSkills): the GitHub-skills generator now resolves outputDir to an absolute path before writing, fixing a regression introduced by #6366 (relative-to-cwd base path) that could write generated skill files to the wrong directory.

  • fix(security): /api/keys/{id}/devices now checks the HTTP method before auth/validation, returning a 405 for non-GET/DELETE verbs instead of a misleading 401/500 (closes a dast-smoke QUERY-method finding).

  • fix(quality): clear the last 2 heavy quality-gate reds on the release tip (cycle pre-flight).

  • fix(mitm): the test suite and CI can never mutate the OS trust store — OMNIROUTE_SKIP_SYSTEM_TRUST=1 is set globally for tests/CI so installCert/uninstallCert/installTproxyCa skip the privileged OS dispatch (#6310; full detail is under the [3.8.45] section below — this branch received it via the parallel-cycle sync-back).

  • fix(api): POST /api/github-skills now Zod-validates its request body; documented the new quality-gate env vars and pinned the merge-integrity GitHub Actions to a commit SHA.

  • fix(skills): generate the missing omni-github-skills registry entry and align the agent-skills catalog-count tests (follow-up to #6186).

  • fix(quality): clear the cycle's 11 net-new ESLint errors and make validate-release-green suppressions-aware.

📝 Maintenance
  • i18n(it): add 118 missing Italian (it) translations (net-additive — no existing keys dropped, valid JSON), improving Italian UI coverage. (#6212 — thanks @serverless83)
  • chore(providers): remove deprecated MiMo V2 model entries from the catalogs (xiaomi-mimo, opencode-go, zenmux-free, audio TTS) — the upstream V2 line is superseded by MiMo V2.5; drops mimo-v2-tts, mimo-v2-pro, mimo-v2-omni, mimo-v2-flash, mimo-v2-flash-free and realigns the provider-catalog tests. (#6248 — thanks @backryun)
  • chore(release): ~50 commits on this branch are v3.8.45 pre-flight/hardening fixes and CI-perf work that landed here via the parallel-cycle sync-back (sync-next-cycle.mjs, Hard Rule #21) after the v3.8.45 git tag was cut, and are already fully documented under the [3.8.45] section below — listed here only so the per-cycle commit-coverage check (npm run release:uncovered) doesn't flag them as gaps. Provider/catalog/UX/backend fixes: #6041, #6078, #6108, #6135, #6148, #6149, #6154, #6158, #6161, #6162, #6163, #6164, #6165, #6170, #6177, #6178, #6181, #6186, #6187, #6191, #6193, #6194, #6195, #6200, #6205, #6208, #6209, #6211, #6213, #6223, #6224, #6225, #6226, #6227, #6228, #6229, #6230, #6235, #6291, #6292. CI/release-pipeline work: #6167, #6203, #6214, #6215, #6218, #6273, #6275, #6283, #6284, #6285, #6300, #6305.
  • chore(release): additional zero-ref release-cycle plumbing on this branch, kept out of release:uncovered on purpose (no #N in the commit subject to cite): opening the v3.8.46 cycle, opening/closing the v3.8.45 cycle, the finalized [3.8.45] CHANGELOG i18n sync-back to 42 mirrors, the v3.8.45 cognitive/cyclomatic and file-size drift rebaselines, ESLint stale-suppression pruning (4,273 → 4,233), and clearing test-masking/docs-all pre-flight reds for v3.8.45.
⚡ Performance & Infrastructure
  • perf(release-green): the pre-flight validator (scripts/quality/validate-release-green.mjs) now runs its 4 slow suites (unit / vitest / integration / pack-artifact) concurrently via Promise.all — pre-flight wall time drops from ~the sum of the suites to ~the slowest one (~30min saved per round; Phase 0 was the nº1 bottleneck of the v3.8.45 release benchmark, 2h54 of 6h34 e2e). Guard: tests/unit/validate-release-green.test.ts ("runs the slow suites CONCURRENTLY"). (#6319)
  • fix(ci): scripts/release/sync-next-cycle.mjs — two defects found live in its first production run (v3.8.45 Phase 5): (1) the git() helper's default 1 MiB maxBuffer crashed with ENOBUFS on git show origin/main:CHANGELOG.md (the CHANGELOG alone is >1 MiB) — widened to 64 MiB; (2) the i18n resync only propagated the [NEXT] (TBD) section, leaving the just-shipped finalized section as "— TBD" in all 42 mirrors — it now also syncs [prevVersion] bounded by the heading below it (new exported pure helper versionAfter). Guards: +5 tests in tests/unit/sync-next-cycle.test.ts (8/8). (#6327)
  • test(ci): concurrency-sensitive flaky tests are quarantined into a serial pass (tests/unit/serial/, --test-concurrency=1, appended to every unit runner incl. sharded variants — the serial pass is sharded too so concurrent shard jobs never self-collide). Initial set: glm-coding-plan-monthly-3580, quota-division-blocks, provider-health-autopilot, combo-health-autopilot — the class behind the ~28min CI wedges/re-runs (two live 1h+ wedges cancelled during this PR's own validation). Discovery + TIA gates track the new glob; systemic root cause (async logger writing after teardown) tracked in #6360. Guard: tests/unit/test-serial-quarantine.test.ts (4). (#6347)
🙌 Contributors

Thanks to everyone whose work landed in v3.8.46:

ContributorPRs / Issues
@2220258345direct commit / report
@abdofallahdirect commit / report
@adentdkdirect commit / report
@aleksesipenkodirect commit / report
@anki1krdirect commit / report
@anungmadirect commit / report
@arpicatodirect commit / report
@backryun#6248
@binsarjrdirect commit / report
@brightfiscalbanddirect commit / report
@chirag127#6501, #6506
@developerjillurdirect commit / report
@dilneiss#6499
@dtybnrjdirect commit / report
@Forcerecondirect commit / report
@hao3039032direct commit / report
@Iammilansoni#6150, #6245
@jmengitdirect commit / report
@jordansilly77-stackdirect commit / report
@JxnLexndirect commit / report
@KooshaPari#6166
@loopyddirect commit / report
@luweiCN#6221
@makcimbxdirect commit / report
@muflifadla38direct commit / report
@newnoldirect commit / report
@ofekbetzaleldirect commit / report
@ohahe52-dotdirect commit / report
@phidinhmanhdirect commit / report
@powellnormadirect commit / report
@qpeybadirect commit / report
@RaviTharumadirect commit / report
@RCrushMedirect commit / report
@rianonehub#6134, #6204
@serverless83#6212
@swingtempodirect commit / report
@tenshiakdirect commit / report
@ThongAccountdirect commit / report
@UnrealAryandirect commit / report
@vinayakkulkarnidirect commit / report
@vittoroliveira-dev#6233
@warelikdirect commit / report
@xxy9468615direct commit / report
@xz-devdirect commit / report
@yanpaing007direct commit / report
@diegosouzapwmaintainer
What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.45...v3.8.46

View originalPermalink
How v3.8.46 went

v3.8.45

Added 9
  • Add Yuanbao (web) as a cookie-session provider with support for Tencent Yuanbao, DeepSeek V3/R1, and Hunyuan models
  • Route the built-in agentrouter through the dynamic Claude-Code wire image while preserving its own registry baseUrl and x-api-key auth
  • Enable bulk-add API keys for Cloudflare Workers AI with per-entry providerSpecificData to avoid shared-object reuse
  • Show effective routing share percentage next to each weight in weighted combos when weights don't sum to 100
  • Add opt-in advanced base-URL override for built-in providers hidden behind an Advanced toggle
  • Add an option to disable session stickiness per-combo or globally to allow round-robin or random combos to rotate connections on every request
Changed 1
  • Rename the status widget's 'Cloud Sync' label to 'Remote Settings Sync'
✨ New Features
  • feat(providers): add Yuanbao (web) as a cookie-session provider (#6196) — yuanbao-web (Tencent Yuanbao, yuanbao.tencent.com) with cookie-only auth (hy_user/hy_token + public agent id), SSE→OpenAI translation incl. reasoning_content, exposing DeepSeek V3/R1 + Hunyuan / Hunyuan-T1. Regression guard: tests/unit/providers-yuanbao-web.test.ts. together-web was deferred (no verifiable web-session endpoint — needs a captured request) and huggingchat-web dropped (the existing huggingchat already is a web-cookie provider). (thanks @chirag127)
  • feat(providers): route the built-in agentrouter through the dynamic Claude-Code wire image (#6056) — a small static allow-set (CC_WIRE_IMAGE_BUILTINS in open-sse/services/ccWireImageBuiltins.ts), consulted by isClaudeCodeCompatible / isClaudeCodeCompatibleProvider / applyFingerprint, makes agentrouter adopt the CC wire-image headers + fingerprint while guarding the CC baseUrl/auth branches so it keeps its own registry baseUrl and x-api-key auth. Regression guard: tests/unit/agentrouter-cc-wire-image.test.ts (asserts the wire image is applied AND agentrouter's baseUrl/auth are preserved). Live WAF-acceptance against agentrouter.org is a VPS validation follow-up (Hard Rule #18).
  • feat(providers): bulk-add API keys for Cloudflare Workers AI (#6174) — cloudflare-ai is removed from the bulk-add exclusion list and the bulk parser gains a 3-field name|accountId|apiKey mode; the bulk route now builds a per-entry providerSpecificData so each key carries its own accountId (fixing the previous shared-object reuse), and both the create + key-validation paths receive it. Regression guard: tests/unit/bulk-api-key-parser-cloudflare.test.ts. (thanks @muflifadla38)
  • feat(dashboard): routing/settings UX clarity (#6147) — (1) weighted combos show the effective routing share % next to each weight when weights don't sum to 100 (WeightTotalBar.tsx); (2) the status widget's user-facing "Cloud Sync" label is renamed to "Remote Settings Sync" (CloudSyncStatus.tsx; internal ids/state untouched); (3) built-in providers gain an opt-in advanced base-URL override (isBaseUrlOverrideEligibleProvider, hidden behind an "Advanced" toggle, reusing the existing providerSpecificData.baseUrl persistence — not globally widened). Regression guard: tests/unit/routing-settings-ux-6147.test.ts.
  • feat(combo): add an option to disable session stickiness, per-combo or globally — round-robin / random combos can rotate to a different connection on every request instead of pinning a whole conversation to one connection by its first-message hash. Resolution precedence per-combo config.disableSessionStickiness → global settings.disableSessionStickiness → default false (preserves the #3825 prompt-cache/504 fix); gates both stickiness call sites in open-sse/services/combo.ts. Exposed as a global toggle (Combo Defaults) and a per-combo Inherit/on/off control. (#6168) Regression guard: tests/unit/combo-disable-session-stickiness.test.ts. (thanks @RCrushMe)
  • feat(docker): add the OMNIROUTE_NO_SUDO env flag for root-less / user-namespaced deployments — the MITM cert-trust command path (resolveSudoSpawn in src/mitm/systemCommands.ts) now strips the leading sudo when the flag is truthy, in addition to the existing root / sudo-missing cases, so the Proxy Agent runs without sudo (the operator trusts the CA manually, e.g. via NODE_EXTRA_CA_CERTS). Argv-array spawn preserved — no shell interpolation (Hard Rule #13). (#6122) Regression guard: tests/unit/mitm-systemCommands-no-sudo.test.ts. (thanks @powellnorma)
  • feat(providers): add Requesty as an OpenAI-compatible gateway provider (BYOK, base https://router.requesty.ai/v1, ~200 free requests/day) — wired through the shared OpenAI-compatible registry with full model passthrough (open-sse/config/providers/registry/requesty/, src/shared/constants/providers/apikey/gateways.ts). (#6120) Regression guard: tests/unit/requesty-provider.test.ts. (thanks @chirag127)
  • feat(dashboard): add configured-only / available-only filters to the Free Provider Rankings page (#6150) — hide providers you haven't configured, or whose connections are all rate-limited / out of quota, via server-side query params (?configuredOnly / ?availableOnly on GET /api/free-provider-rankings) backed by a testable lib helper reusing the in-process connection state (no Redis). Both filters default off, so the default view is unchanged; this supersedes the earlier client-side "Configured Only" toggle (#6245) with an available-only dimension and unit-tested logic. Regression guard: tests/unit/freeProviderRankings-filters.test.ts.
  • feat(rankings): add a 'Configured Only' filter to the Free Provider Rankings page, so the table can be narrowed to just the providers you have configured connections for (with an empty-state hint when none are configured). New en.json keys and a pure filter helper covered by tests/unit/free-provider-rankings-configured-filter.test.ts. (#6245, closes #6150 — thanks @Iammilansoni)
🔧 Bug Fixes
  • fix(mitm): the test suite and CI can never mutate the OS trust store again — OMNIROUTE_SKIP_SYSTEM_TRUST=1 (set by the global test setup and all CI workflows) makes installCert/uninstallCert/installTproxyCa skip the privileged OS dispatch while preserving the #4546 environment-skip contract. Root cause of the self-hosted runner incident: a cert-flow integration test installed a 105-byte fake PEM into /usr/local/share/ca-certificates, breaking ALL system TLS on the VM. Regression guard: tests/unit/system-trust-test-guard.test.ts. (#6310)

  • fix(security): /api/keys/{id}/devices answers a clean method-first 405 for undocumented HTTP methods (e.g. the new QUERY) via a dedicated http-method-guard rule — the auth layer was answering 401 first, failing schemathesis's unsupported-methods check. Same pattern as the v3.8.44 TRACE fix. Regression guard: tests/unit/dast-method-not-allowed.test.ts.

  • fix(combo): the #6216 empty-stream failover is restricted to truly empty bodies (zero bytes — the Gemini HTTP-200-empty case), restoring the #3399/#3685 pass-through contracts for [DONE]-terminated empty streams and incomplete Claude lifecycles. New guard: #5976 truly EMPTY streaming body → invalid for combo failover (87/87 across both suites).

  • fix(combo): 5 streaming-path fixes — locked-stream 500, error-frame-only-if-no-content, Gemini MALFORMED_RESPONSE→content_filter failover, correlationId substring search, per-model-500 lockout skip + request-logger UI detail. Maintainer follow-up: releaseQualityClone cancels the abandoned quality-check tee branch (per-request memory) + regression test. (#6216 — thanks @hartmark)

  • fix(skills): generate the missing omni-github-skills registry entry (the #6186 catalog addition never ran the generator — 8 integration assertions split between old/new counts) and align the agent-skills catalog counts across integration + unit suites (43 = 23 API + 20 CLI; 44 with config).

  • fix(a2a): finish the #6186 catalog-count update — listCapabilities metadata reported coverage.api.total: 22 (type literal + value) and SkillCoverageSchema pinned z.literal(22), so the schema would REJECT the correct runtime value with 23 API skills. All three aligned to 23.

  • fix(github-skills): add a missing import, unit tests and a settings JSON-parse fix for the GitHub agent-skill discovery/import flow. (#6186 — thanks @Moseyuh333)

  • fix(api): POST /api/github-skills validates its body with a Zod schema (validateBody) instead of blind request.json() destructuring — a non-array targets would crash .map. Regression guard: tests/unit/github-skills-route-validation.test.ts.

  • fix(docker): add id= to the BuildKit cache mounts so strict builders (e.g. buildkitd with strict frontend parsing) accept the Dockerfile. (#6291 — thanks @karimalsalah)

  • fix(oauth): register zed in the OAuth PROVIDERS map (fixes "Unknown provider" on the Zed sign-in flow) (#6078 — thanks @anki1kr), and align zed in OAUTH_PROVIDER_IDS + the config enum after the merge.

  • fix(doubao-web): switch the Doubao web provider to the Dola global endpoint. (#6235 — thanks @backryun)

  • fix(doctor): resolve two false-positive WARNs in the doctor diagnostics (#6163, closes #6162 — thanks @arssnndr)

  • fix(providers): refresh the GitHub Copilot model catalog to the current upstream set. (#6154 — thanks @backryun)

  • fix(providers): correct the Kiro model catalog to real upstream ids — fabricated claude-opus-4.7/claude-sonnet-4.6 entries removed, real claude-sonnet-5/claude-sonnet-4.5/claude-haiku-4.5 kept. (#6170)

  • feat(sse): surface Kiro adaptive-thinking reasoning frames as reasoning_content in the OpenAI-shaped stream. (#6213 — thanks @VXNCXNX)

  • fix(cli): use OMNIROUTE_SERVER_HOST instead of the POSIX auto-set HOSTNAME for the bind address (fixes wrong bind on POSIX shells that export HOSTNAME). (#6195, closes #6194 — thanks @Theadd)

  • feat(provider): add Claude 5 Sonnet to the Claude Web provider catalog. (#6209, closes #6200 — thanks @Iammilansoni)

  • fix(providers): add nvidia to PROVIDER_TOOL_LIMITS (1536) to prevent silent tool-list truncation. (#6177 — thanks @LuisAlejandroVega)

  • fix(translator): strip the reasoning param for nvidia z-ai/glm-5.2 (upstream 400s on it). (#6181 — thanks @kanztu)

  • fix(dashboard): providers page gains a data-timeout guard and the live-WS standalone wiring (no more indefinite spinner when the data fetch stalls). (#6211)

  • fix(sse): surface the ChatGPT-web image silent-drop as an accurate error instead of an empty success. (#6208)

  • fix(cline): force upstream streaming for Cline/ClinePass (streaming-only API) — non-stream client requests are served from the buffered SSE. (#6165)

  • fix(dashboard): remove the always-on Auto-Routing (combo) banner from the home page — it did not reflect live routing state and reappeared on every fresh browser. Replacement guard: tests/unit/home-no-autorouting-banner.test.ts. (#6164)

  • fix(dashboard): stop a model-test error from freezing the page (React #31 object-as-child toast) — errors go through extractApiErrorMessage. (#6161)

  • fix(oauth): extract the keychain-import-only guard to its own module, restoring the oauth file-size freeze. (#6158)

  • fix(sse): strip zero-width markers from streamed tool-call arguments — a follow-up to #5857. That PR removed injected zero-width joiners (U+200D) from streamed assistant text/reasoning but deliberately left tool-call argument JSON byte-exact. The request-side obfuscation (open-sse/services/claudeCodeObfuscation.ts) injects ZWJ into agent words — including the temp path inside the Bash tool description — and Claude models copy that verbatim into generated commands, which are delivered as tool-call arguments rather than assistant text. As a result the ZWJ survived and corrupted code blocks (e.g. a temp path rendered with an invisible joiner). Now open-sse/handlers/responseSanitizer.ts strips zero-width code points from tool-call argument strings at every emit site (OpenAI non-stream/stream chat tool_calls + legacy function_call, native Responses function_call items, the OpenAI→Responses conversion, and the native Responses streaming response.function_call_arguments.delta/.done events). Only zero-width code points are removed; JSON structure and all other bytes stay identical (no parse/restringify), so normal arguments remain byte-exact. Regression guard: 6 new cases in tests/unit/response-sanitizer.test.ts (suite 50/50).

  • fix(nodejs): the default app log path now resolves under DATA_DIR (~/.omniroute/logs/application/app.log) instead of process.cwd() (#6197) — the globally-installed CLI runs from an arbitrary working directory, so anchoring the default to cwd made file logging silently write to (or no-op under) an unrelated directory, contradicting the documented .env.example default. getAppLogFilePath() now computes the default lazily via the pure resolveDataDir() resolver (honours a per-process DATA_DIR, no directory-creation side effect); an explicit APP_LOG_FILE_PATH still wins. Regression guard: tests/unit/logenv-datadir-path-6197.test.ts (3). (root cause independently diagnosed by @subhansh-dev in #6298 — thanks!)

  • fix(docker): AgentBridge/startMitm no longer aborts in containers/headless when the Antigravity-default DNS step can't write /etc/hosts (#6127), and the privileged command's stderr now reaches app.log instead of only a bare exit code hitting the toast (#6198). The default DNS step (addDNSEntry) was called unguarded while cert install and the two sibling DNS steps were each best-effort — in the runtime Docker image (USER node, no sudo, read-only /etc/hosts) it threw Command failed with code 1 out of startMitmInternal and killed the whole start, discarding the stderr. The three DNS steps are extracted into a best-effort provisionDnsEntries() where each failure is logged with the full err (stderr included, folded in by systemCommands.ts) and never aborts the start. Regression guard: tests/unit/mitm-dns-graceful-degrade-6127.test.ts (4).

  • fix(providers): copilot-m365-web now supports the M365 Education "Starter / OfficeWebIncludedCopilot" tier and no longer returns an empty content:null stream (#6210). Two gaps: (1) buildWsUrl() hardcoded the individual-consumer scenario (OfficeWebPaidConsumerCopilot, isEdu=false) — the EDU tier is now opt-in via providerSpecificData.tier="edu", emitting scenario=OfficeWebIncludedCopilot/isEdu=true (the individual path is unchanged); (2) the EDU/GPT-5.5 path streams deltas via arguments[0].writeAtCursor (incremental) instead of only messages[].text (accumulated snapshots), which the parser dropped — a new accumulateBotContent() folds both formats, with type:2 item.result.message as a last-resort fallback. Regression guard: tests/unit/copilot-m365-edu-writeatcursor-6210.test.ts (10). (thanks @qpeyba)

  • fix(providers): GitLab Duo executor now feeds tool results back into the prompt instead of looping (#6220) — buildPrompt() branched only on system/user and took userParts.at(-1), silently dropping the assistant{tool_calls} + tool{result} turns the client appended, so the reconstructed prompt was byte-identical to turn 1 and the model re-emitted the same <tool> call forever. When a tool exchange is present the full conversation is now serialized, folding each tool result back keyed by its tool_call_id; simple conversations keep the legacy shape. Complements the tool_call emission from #6051 (the kilo-duplicate label was a false positive — different, sequential defect). Regression guard: tests/unit/gitlab-tool-result-feedback-6220.test.ts (4).

  • fix(providers): opencode-go/opencode-zen can now synthesize the OpenCode CLI identity headers Cloudflare requires on VPS egress (#5997) — on a datacenter VPS, opencode.ai/zen/go/v1/chat/completions 403s (HTML challenge) requests lacking CLI identity, while the reporter's control curl proved that User-Agent: opencode-cli/1.0.0 + x-opencode-client: cli + x-opencode-project: default + fresh request/session UUIDs succeed. Opt-in via OPENCODE_SYNTHESIZE_CLI_HEADERS=true (values overridable via OPENCODE_GO_USER_AGENT/OPENCODE_USER_AGENT/OPENCODE_CLIENT/OPENCODE_PROJECT); it fills only headers the client did not already send. Kept off by default — the forward-only path is deliberate (fabricating a wrong value risks upstream rejection; a prior dedup regressed with opencode/local), so this replaces the fragile local header-injection shim without changing default behavior. Regression guard: tests/unit/opencode-cli-headers-synthesis-5997.test.ts (6). (thanks @aleksesipenko)

  • fix(resilience): sticky session affinity now evicts and fails over to another account when the pinned account is exhausted/unavailable (#6219)

  • fix(sse): Responses API passthrough now drops internal commentary-phase output before forwarding to clients (gated by RESPONSES_PASSTHROUGH_DROP_COMMENTARY, default on) (#6199)

  • fix(proxy): stop the v3.8.44 proxy regression that leaked the real IP and disabled healthy proxies (#6246). Two coupled defects from the new health scheduler: (1) IP leak — when a proxy assigned to a connection was marked inactive, resolution fell through to a direct egress instead of blocking, exposing the operator's real IP; (2) over-deactivation — the sweep flipped a proxy to inactive on the first failed probe and counted our own 5s timeout / a probe-target 5xx as the proxy's fault, so healthy paid proxies vanished from egress selection ("my proxies are not being used anymore"). Fix: the sweep decision is extracted into a pure, network-free decideProxyHealthAction (src/lib/proxyHealth/decision.ts) — by default the health check now only counts/logs and never downgrades status (a proxy is downgraded/removed only with PROXY_AUTO_REMOVE=true, after PROXY_AUTO_REMOVE_AFTER consecutive conclusive failures); probes are classified tri-state so an inconclusive result (our timeout, or a 5xx from the probe target) never penalizes the proxy, and the probe timeout is raised 5s→15s. Separately, safeResolveProxy now fails closed via the existing policy: a connection whose assigned proxy is dead is blocked instead of leaking direct (hasBlockingProxyAssignment), honoring the explicit proxy off toggles and the PROXY_FAIL_OPEN=true opt-out. Existing proxies stuck inactive by the old behavior need a one-time manual re-activate (the operator owns proxy status). Regression guards: tests/unit/proxy-health-decide-action-6246.test.ts, tests/unit/proxy-assigned-unavailable-6246.test.ts.

  • fix(proxy): make "Test All" read-only and add bulk enable/disable (#6246). Complements the core fail-closed / scheduler fix (#6296) with the two remaining reporter asks. (1) The "Test All" button (POST /api/settings/proxies/auto-test) used to flip a proxy to inactive on a failed reachability probe; since the egress selector excludes inactive proxies, a flaky probe (an unreachable httpbin.org, a proxy that blocks HEAD, or a slow paid proxy) silently disabled every proxy that failed — "Test All" is now read-only by default (only the operator sets a proxy active/inactive; opt back into the legacy test-and-set with PROXY_HEALTH_AUTO_DEACTIVATE=true). (2) Adds a bulk enable/disable proxies endpoint + toolbar action (POST /api/settings/proxies/batch-activate) so an operator can re-activate proxies in one click. Regression guard: tests/unit/proxy-health-6246.test.ts. (thanks @tenshiak)

  • chatcore (tools): stop the default 128-tool cap from silently dropping opencode's task/MCP tools. opencode (used as an MCP/agent host) sends a large tool list; when it exceeds the speculative MAX_TOOLS_LIMIT (128) default, truncateToolList did a blind tools.slice(0, 128), dropping every tool past index 128 — including opencode's built-in task tool (subagent launch) and many MCP tools, so models routed through OmniRoute could no longer spawn subagents or reach part of their tools. The cap exists to avoid upstream 400s for providers with real hard limits (e.g. grok-cli 200), so it is kept for those: detection of the opencode client (isOpencodeClient — any x-opencode-* header, or opencode in the user-agent) now only bypasses the speculative 128 default, never a known provider ceiling. Precedence is explicit — a proactive/detected provider limit always truncates (even for opencode); otherwise opencode forwards its full tool list; otherwise the unchanged 128 default applies to every other client. Refactors getEffectiveToolLimit into getKnownToolLimit(provider) ?? DEFAULT_LIMIT (byte-identical for existing callers) and fixes a cosmetic debug-log that reported the truncated count instead of the original. Regression guard: tests/unit/tool-limit-detector.test.ts. (#6193 — thanks @DKotsyuba)

  • fix(mitm): the macOS MITM-cert install check now matches the system keychain again. security find-certificate -a -Z prints the SHA-1 as a colon-less hex string, but the installed-check compared it against getCertFingerprint()'s colon-separated form, so the substring match never hit — the cert was reported as not-installed and re-prompted for the sudo install on every run. Fingerprints are now normalized (colons stripped, upper-cased) on both sides via the extracted macCertOutputHasFingerprint helper. Regression guard: tests/unit/mitm-cert-mac-fingerprint.test.ts. (#6204, closes #6134 — thanks @rianonehub)

  • fix(api): /v1/messages/count_tokens now counts tool_use, tool_result and thinking content blocks (and array-form system prompts) in the local-estimation path, instead of only text. Real agentic conversations keep ~95% of their tokens inside tool results; the previous estimate returned near-zero for them, which silently broke Claude Code's auto-compaction (context grew past the window with no compaction until the upstream API rejected the request). The real provider-side count path is unchanged. Regression guard: tests/unit/messages-count-tokens-route.test.ts. (#6221 — thanks @luweiCN)

  • fix(antigravity): strip a trailing assistant prefill turn for Vertex Claude models to avoid upstream 400s (#6114). Regression guard: tests/unit/antigravity-claude-prefill-strip.test.ts. (thanks @anki1kr)

  • fix(security): the mutable cloud-agent routes (/api/cloud/credentials/update, /api/cloud/models/alias) now require management auth instead of being treated as public. They were classified as public API routes, so a request without management credentials could update stored cloud-agent credentials and model aliases. They are removed from the public-route set, classified as management routes in the authz pipeline, and gated by requireManagementAuth; cloud read/auth routes stay public. Regression guards: tests/unit/cloud-write-auth.test.ts, tests/unit/authz/classify.test.ts, tests/unit/public-api-routes.test.ts. (#6233 — thanks @vittoroliveira-dev)

  • refactor(dashboard): extract the onboarding-wizard "Open provider details" link target into a pure, unit-tested buildProviderDetailsHref(connection) helper. The wizard already routes by connection.id (the node UUID) rather than the provider category slug (#6144/#6145); this hardens that behavior behind a tested helper that guards a missing id/connection. Regression guard: tests/unit/provider-onboarding-href.test.ts. (#6166 — thanks @KooshaPari)

  • fix(api): relay worker now binds the SSRF guard to a stable const name so minified standalone (Docker) builds resolve it (#6149) — the Vercel/Deno relay generators embedded the shared resolveRelayTarget guard as a bare ${fn.toString()} declaration while the worker body called the hardcoded literal name; SWC minification mangled the source function's name, so the deployed worker defined <mangled> but still called resolveRelayTargetReferenceError. Both templates now emit const resolveRelayTarget = ${fn.toString()}; (the const name is a template literal, immune to minification). Regression guard: tests/unit/relay-minified-fn-6149.test.ts (4). (thanks @SeaXen)

  • fix(providers): refresh the stale NVIDIA NIM model registry — drop EOL z-ai/glm-5.1, add z-ai/glm-5.2 and nvidia/nemotron-3-ultra-550b-a55b (#6108). Regression guard: tests/unit/nvidia-nim-registry-6108.test.ts. (thanks @andrea-kingautomation)

  • fix(backend): GPT-family (codex) models now report a distinct max_input_tokens (272000) below their 400K context_length via an optional maxInputTokens on RegistryModel, so coding agents auto-compact correctly instead of overflowing the real input cap (#6191). Regression guard: tests/unit/gpt-max-input-tokens-6191.test.ts. (thanks @luweiCN)

  • fix(backend): call logs now record a reasoning source/char-count (migration 116, reasoning_source/reasoning_chars) for models that emit reasoning_content/<think> but report zero reasoning tokens in usage, so tokens_reasoning no longer silently under-represents reasoning — cost math is unchanged (the priced tokens_reasoning stays usage-derived) (#6187). Regression guard: tests/unit/reasoning-token-source-6187.test.ts. (thanks @andrea-kingautomation)

  • fix(auth): a stale/changed STORAGE_ENCRYPTION_KEY now surfaces as a clear 424 storage_encryption_stale ("re-enter the API key") instead of a misleading "Auth failed: 401" — the connection's ciphertext failed to decrypt and was coerced to an empty Bearer, hiding the real cause (#6148). Regression guard: tests/unit/decrypt-stale-key-hint-6148.test.ts. (thanks @chirag127)

  • fix(backend): memory injection now keeps the injected system message first for providers that require it (via a PROVIDERS_SYSTEM_MUST_BE_FIRST capability), instead of the cache-safe mid-array splice that made strict providers reject the request with a 400 (#6135). Regression guard: tests/unit/memory-system-first-6135.test.ts.

  • fix(services): 9Router embed panel no longer 404s (optional catch-all route) and the supervisor probes the port before spawning to avoid raw EADDRINUSE (#6205). Regression guards: tests/unit/ninerouter-embed-port-6205.test.ts, tests/unit/services/ServiceSupervisor.test.ts. (thanks @jonlwheat2-gif)

  • fix(mcp): forward the MCP request extra context through static tool loops so stdio callers keep their scope/identity (#6178)

⚡ Performance & Infrastructure
  • perf(test): test-suite loader quick wins (#6214) — the 19 test scripts switch --import tsx--import tsx/esm (the repo is pure ESM; the unused CJS hook cost ~1.3s per test process × 2,462 processes — CI fast-path unit shards dropped 14.8→7.5 min, −49%), tsx bumped to ^4.23.0 (tsx#809 startup-regression fix), 37 orphan .test.mjs files (224 cases) recovered into the canonical glob (they matched no runner and never ran in any CI job; check:test-discovery now scans .mjs too), and ci.yml/quality.yml unit jobs now call the canonical npm script test:unit:ci:shard (single source of truth — closes two silent drifts: missing setupPolyfill import in CI and memory/+usage/ dirs absent from the fast-path glob). tests/unit/dashboard/** keeps the full tsx hook in its own invocation (@lobehub/icons es/ build internally require()s ESM-syntax files).
  • ci: heavy-pipeline dedup (#6215) — the release-PR pipeline ran the unit suite 4× per sync (95 jobs, 208 machine-min; the v3.8.44 cycle fired 123 such runs, 88 cancelled). Now: Node 24/26 compat matrices move to a daily nightly-compat.yml (−28%/run; resolves the active release branch, opens a tracking issue on failure), coverage is collected inside the unit shards themselves via c8/NODE_V8_COVERAGE (−18%/run; the Coverage Shard ×8 matrix is gone — nodejs/node's own CI pattern), the ~40-job per-language i18n matrix becomes 1 job (the account has 20 concurrent-job slots total), and heavy jobs skip draft PRs — paired with /generate-release now opening the living release PR as draft (flipped ready at the new Phase 0a.0a), killing the per-merge churn for the whole cycle. Validated by a full workflow_dispatch of the new pipeline: 35 jobs, 0 failures, 23 min, merged coverage 80.16% (> ratchet baseline).
  • feat(quality): no-new-warnings per PR (#6218) — native ESLint bulk suppressions (≥9.24) freeze the pre-existing debt (476 files / 4,273 violations in config/quality/eslint-suppressions.json); npm run lint, lint-staged (pre-commit) and a new fork-aware lint-guard job in quality.yml all run suppressions-aware, so a NEW warning goes red in the PR that introduces it instead of accruing invisibly (+41/+88 per cycle) and being blind-rebaselined at release. 3 warn rules promoted to error in src/** (react-hooks/exhaustive-deps, @next/next/no-img-element, import/no-anonymous-default-export); collect-metrics measures under the frozen baseline (ratchet metric = net-NEW debt; baseline tightened 4,279→0 in-PR per require-tighten); fork PRs run report-only (contributors are never blocked — the maintainer campaigns fix via co-authorship). Baseline stock shrinks via --prune-suppressions at release reconciliation.
  • ci: test jobs no longer wait on the Build gate (#6275) — test-unit×8, vitest, integration×2 and security declared needs: build but never download the next-build artifact; they now start at minute 0 (needs: changes, same if as Build), cutting ~15–20 min of wall-clock per heavy run. e2e/package-artifact/electron-smoke keep needs: build (they consume the artifact for real).
  • ci(build): the ci.yml Build job compiles Next.js with Turbopack (OMNIROUTE_USE_TURBOPACK=1) (#6273) — Build job 20 min → 6 min 59 s (~2.9×) on ubuntu-latest; the webpack actions/cache step is removed. Validated end-to-end pre-merge via gh workflow run ci.yml --ref <branch>.
  • feat(build): Turbopack becomes the default bundler for next build and next dev (#6283) — build-next-isolated.mjs, run-next.mjs and the playwright-runner default to Turbopack; OMNIROUTE_USE_TURBOPACK=0 is the explicit webpack escape hatch. nightly-compat.yml/npm-publish.yml inherit the default. Regression guard: tests/unit/build-bundler-default-turbopack.test.ts.
  • feat(docker): the Docker image builds with Turbopack (ENV OMNIROUTE_USE_TURBOPACK=1) (#6285) — the v3.8.27 ImportTracer panic ("unreachable: there must be a path to a root") does not reproduce on Next 16.2.9: amd64 (659 s) and arm64 (qemu) build clean, 0 panics, smoke health 200.
  • ci: opt-in self-hosted VPS runners for the release window (#6284) — scripts/vps/release-runner-up.sh/down.sh manage the runner VM, and build/test-unit/vitest pick a dynamic runs-on gated by vars.USE_VPS_RUNNER == 'true' and own-origin (fork PRs never reach self-hosted runners). Wired into /generate-release (VM up at Phase 1, mandatory down at Phase 3).
📝 Maintenance
  • quality(release-green): full pre-flight hardening for this release — the cycle's 11 net-new ESLint errors typed/fixed and validate-release-green made suppressions-aware with per-gate logs (_artifacts/release-green/) and a --hermetic mode; test-masking allowlist entries for the cycle's verified-legitimate assert reductions; stale ESLint suppressions pruned (4,273 → 4,233); the 7 net-new as any casts from #6292 typed; githubSkillTools MCP errors routed through sanitizeErrorMessage(); combo-provider-cooldown-sibling added to the Stryker tap set; executors/env docs count fixes.
  • ci(quality): merge-integrity fast-gates per PR — check:changelog-integrity (no base CHANGELOG bullet may vanish in the merge result — the auto-resolve "CHANGELOG-eat" pattern) and check:agent-skills-sync (generated SKILL.md ≡ catalog), blocking for own-origin branches and report-only for forks (Princípio Zero). (#6300)
  • ci(vps): hermetic nightly-release-green pre-flight on the dedicated omni-release self-hosted runner (dynamic runs-on, clean env); e2e/integration/electron stay on hosted runners (per-VM port collision + concurrent artifact-download limits documented in the PR). (#6305)
  • chore(quality): v3.8.45 cycle-close drift rebaselines — file-size (13 files grown by merged cycle PRs), cognitive 867→877, cyclomatic 2028→2035, kiro-translator test debt from #6213; all with dated justification keys.
  • docs(architecture): sync stale DB-layer counts (45+/55 → 95+/110+) in REPOSITORY_MAP, the db-schema diagram and llm.txt (+42 i18n mirrors). (#6167)
  • chore(release): parallel-cycle flow — sync-next-cycle.mjs + Hard Rule #21 semantics (#6203); v3.8.45 development cycle opened.
  • i18n(it): add 118 missing Italian (it) translations (net-additive — no existing keys dropped, valid JSON), improving Italian UI coverage. (#6212 — thanks @serverless83)
  • chore(providers): remove deprecated MiMo V2 model entries from the catalogs (xiaomi-mimo, opencode-go, zenmux-free, audio TTS) — the upstream V2 line is superseded by MiMo V2.5; drops mimo-v2-tts, mimo-v2-pro, mimo-v2-omni, mimo-v2-flash, mimo-v2-flash-free and realigns the provider-catalog tests. (#6248 — thanks @backryun)
🙌 Contributors

Thanks to everyone whose work landed in v3.8.45:

ContributorPRs / Issues
@aleksesipenkodirect commit / report
@andrea-kingautomationdirect commit / report
@anki1kr#6078
@arssnndr#6162, #6163
@backryun#6154, #6235, #6248
@chirag127direct commit / report
@DKotsyuba#6193
@hartmark#6216
@Iammilansoni#6150, #6200, #6209, #6245
@jonlwheat2-gifdirect commit / report
@kanztu#6181
@karimalsalah#6291
@KooshaPari#6166
@LuisAlejandroVega#6177
@luweiCN#6221
@Moseyuh333#6186
@muflifadla38direct commit / report
@powellnormadirect commit / report
@qpeybadirect commit / report
@RCrushMedirect commit / report
@rianonehub#6134, #6204
@SeaXendirect commit / report
@serverless83#6212
@subhansh-dev#6298 (diagnosis, landed via #6234)
@tenshiakdirect commit / report
@Theadd#6194, #6195
@vittoroliveira-dev#6233
@VXNCXNX#6213
@diegosouzapwmaintainer
What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.44...v3.8.45

View originalPermalink
How v3.8.45 went

v3.8.44

Added 12
  • Add throttling for upstream quota fetches on the per-request preflight path via a global min-interval gate, configurable via OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS (default 250ms, clamped 0..5000)
  • Add per-request Auto-Combo controls via X-OmniRoute-Mode and X-OmniRoute-Budget headers to steer scoring and set a hard per-request USD cost ceiling
  • Add the Kenari OpenAI-compatible gateway (BYOK) as a provider
  • Add claude-sonnet-5 to the Antigravity model catalog
  • Add /v1/ocr endpoint (Mistral OCR), an OCR provider category, and Mistral moderation support
  • Add discoveryResults DB module with CRUD operations and persist provider-discovery findings through the discovery_results table
Changed 1
  • Refresh The Old LLM (Free) model catalog with current free tier models (GPT-5/5.1/5.2/5.3/5.4, o3/o4-mini, Gemini 3 Pro / 2.5 Pro / 2.0 Flash / 1.5 Flash, Claude 4.6 Opus/Sonnet & 4.5 Haiku, GPT-4o, Grok 4, DeepSeek V3/R1, Sonar Pro) while keeping legacy alias IDs
Fixed 1
  • Fix mapModel() to pass known upstream IDs through unchanged so Gemini/o-series/Grok/DeepSeek/Sonar models no longer collapse onto GPT_5_4
✨ New Features
  • feat(resilience): throttle upstream quota fetches on the per-request preflight path (#6009) — a new global min-interval gate (open-sse/services/quotaFetchThrottle.ts) spaces the actual network calls made by the Codex quota fetcher so that many accounts on one IP no longer fetch quota in the same second (which, per router-for-me/CLIProxyAPI#2385, can get a Codex OAuth token revoked). Complements the existing bulk-sync spacing (PROVIDER_LIMITS_SYNC_SPACING_MS) which already serialized the periodic provider-limits sync — this covers the concurrent combo/preflight path it didn't. Cache hits are never delayed; fail-open (only ever awaits a timer). Configurable via OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS (default 250ms, clamped 0..5000; 0 disables). Regression guard: tests/unit/quota-fetch-throttle-6009.test.ts (5). (thanks @powellnorma)
  • feat(autoCombo): add per-request Auto-Combo controls via two headers (#6024 / #6025 / #6023) — X-OmniRoute-Mode steers an auto combo's scoring for a single request (friendly presets fast/balanced/quality/cheap/reliable/offline or a raw mode-pack name; balanced forces the default weights), and X-OmniRoute-Budget sets a hard per-request USD cost ceiling. Both override the combo's stored config only for the request that carries them; unknown/garbage values are ignored so the saved config is preserved. The resolvers are pure (open-sse/services/autoCombo/requestControls.ts) and feed the engine's existing config.modePack / config.budgetCap inputs — no engine changes. Regression guard: tests/unit/auto-combo-request-controls-6024.test.ts (5). (thanks @chirag127)
  • feat(providers): add the Kenari OpenAI-compatible gateway (BYOK). Regression guard: tests/unit/kenari.test.ts. (thanks @doedja)
  • feat(models): add claude-sonnet-5 to the Antigravity model catalog (alias mapping in antigravityModelAliases.ts) (#6103). Regression guard: tests/unit/antigravity-model-aliases.test.ts. (thanks @anki1kr)
  • feat(api): add /v1/ocr endpoint (Mistral OCR), an OCR provider category, and Mistral moderation support. (#5950) (thanks @waguriagentic)
  • Discovery tool (Phase 2): add the discoveryResults DB module (CRUD over the discovery_results table, migration 074) and wire the opt-in provider-discovery service to persist and read findings through it (persistDiscoveryResult, getDiscoveryResults, getDiscoveryResultById, markVerified, deleteDiscoveryResult) with (provider, method, endpoint) upsert de-duplication. Adds the /api/discovery/* HTTP surface — GET /results, GET|DELETE /results/:id, POST /scan, POST /verify/:id — under strict loopback-only authorization (/api/discovery/ is in LOCAL_ONLY_API_PREFIXES and is NOT manage-scope-bypassable, so the scan route's outbound probes can never be reached from a tunnel/remote origin). Adds a dashboard UI tab (Tools → Discovery, /dashboard/discovery) to run scans and review, verify, or delete findings. The service stays opt-in / default-off. (#5939)
  • feat(api): expose a read-only provider plugin manifest at GET /api/v1/provider-plugin-manifest for sidecar/relay discovery. (#6001) (thanks @KooshaPari)
  • feat(sidecar): advertise the provider manifest URL to Bifrost/CLIProxyAPI via the X-OmniRoute-Provider-Manifest-Url header (OMNIROUTE_PROVIDER_MANIFEST_URL). (#6007) (thanks @KooshaPari)
  • feat(autoCombo): add a latency/speed-optimized routing mode (shared rankBySpeed scoring core) plus the omniroute_pick_fastest_model MCP tool. (#6011) (thanks @KooshaPari)
  • feat(providers): refresh The Old LLM (Free) model catalog (#5181) — seed the current free /api/chatgpt tier (GPT-5/5.1/5.2/5.3/5.4, o3/o4-mini, Gemini 3 Pro / 2.5 Pro / 2.0 Flash / 1.5 Flash, Claude 4.6 Opus/Sonnet & 4.5 Haiku, GPT-4o, Grok 4, DeepSeek V3/R1, Sonar Pro) while keeping the legacy alias IDs for saved-preference compatibility. Also fixes a latent routing bug: mapModel() now passes known upstream IDs through unchanged, so Gemini/o-series/Grok/DeepSeek/Sonar models no longer silently collapse onto GPT_5_4. Regression guard: tests/unit/theoldllm-model-refresh-5181.test.ts. (thanks @WslzGmzs)
  • feat(resilience): surface Codex banked reset credits per connected account (#5199) — the Codex quota parsers (buildCodexUsageQuotas, parseCodexUsageResponse) now additively read rate_limit_reset_credits.available_count (+ optional rate_limit_reached_type) from the /wham/usage payload OmniRoute already fetches, and the provider-limits dashboard renders a "Banked Reset Credits" row when a positive count is present. Display-only and fail-open — the field is eligibility-gated, so accounts without it are unaffected (parsers never throw on absent/garbage shapes); redemption (an unofficial mutating endpoint) is intentionally out of scope. Regression guard: tests/unit/codex-banked-reset-credits-5199.test.ts (8). (thanks @ofekbetzalel)
  • feat(providers): add sign-up geo-restriction notices for SenseNova and StepFun (#5462) — the provider add-form now warns that SenseNova's console appears to require a Chinese (+86) phone number with no documented international path, and that StepFun's default endpoint is its China platform while a global StepFun Open Platform (platform.stepfun.ai, operated by Sparkling AI Pte. Ltd., Singapore) with email/Google/Discord login exists for international users. Informational notice only — neither provider is disabled. Regression guard: tests/unit/regional-provider-cn-notices-5462.test.ts. (thanks @chirag127)
  • feat(usage): add on-demand period-scoped usage-data reset (Settings → System Storage) with a purge API and time-window selector. (#5831)
  • feat(claude-code): add an opt-in auto-permission classifier compat mode (off/auto/always) for Claude Code, toggleable from the CLI Code settings. (#5810)
  • feat(providers): add optional client-identity header profiles for compatible nodes — preset User-Agent/fingerprint headers (e.g. matching a known CLI) merged into the existing customHeaders field. (#5812)
  • feat(build): add a backend-only fast build mode (scripts/build/build-next-isolated.mjs + backendOnlyPages.mjs) that skips compiling the dashboard frontend pages, cutting local/CI build time for backend-only changes. (#6119 — thanks @artickc)
  • feat(minimax): extract MiniMax M3's raw <think>...</think> leakage into reasoning_content on the 8 OpenAI-format provider tiers, leaving the Claude-format minimax/minimax-cn tiers untouched (they already report reasoning correctly). (#6073 — thanks @KooshaPari)
  • feat(services): promote Bifrost (@maximhq/bifrost — Go AI-gateway) from an env-only relay sidecar to a first-class embedded/supervised service, matching the existing cliproxy/9router model — installer, bootstrap SERVICES[] entry, migration 113 DB seed, 7 lifecycle API routes under /api/services/bifrost/ (loopback-only), a dashboard tab, and relay auto-wiring that defaults BIFROST_BASE_URL to the supervised port when running. Implements item #2 of #5670; the broader RouterBackend contract (items #1, #3-#5) stays out of scope. (#5817, part of #5670)
  • feat(services): add Mux (coder/mux — local agent-orchestration daemon) as a fourth-tier embedded service on the existing ServiceSupervisor framework — npm-based installer, bootstrap.ts registration, migration 113 DB seed, 7 lifecycle API routes under /api/services/mux/ (loopback-only, defense-in-depth bind to 127.0.0.1), and a dashboard tab reusing the shared service-management components. (#6034)
  • feat(xai): surface Grok/xAI usage on the quota dashboard via local usageHistory aggregation (getXaiUsage) — since xAI exposes no per-account quota API, this sums tokens routed to the connection from usage_history and reports them as a cumulative, uncapped quota, mirroring the existing Xiaomi MiMo self-track pattern. (#5806)
  • feat(minimax): extract MiniMax M3's raw <think>...</think> tags into a separate reasoning_content field on the 8 provider tiers that register M3 with format:"openai" (trae, huggingchat, bazaarlink, ollama-cloud, opencode, cline, opencode-zen, codebuddy-cn) — previously the thinking text leaked directly into content. Reuses the existing extractThinkingFromContent primitive, extending its allowlist with a minimax-m3-only pattern; the two direct minimax/minimax-cn tiers are untouched since they already surface reasoning natively over Anthropic's Messages format. (Inspired by 9router#2231.) (#6050 — thanks @KooshaPari)
  • feat(i18n): auto-detect the browser language on first visit — a pure detectBrowserLocale() matcher (exact match, zh-HK/zh-MO folded to zh-TW, language-prefix match, else null) plus a client-only LocaleAutoDetect component mounted once in the root layout. When no locale cookie is set yet, it reads navigator.languages, computes a match against the supported locales, and persists it via the same cookie/localStorage writer LanguageSelector already used (extracted to shared/lib/persistLocale.ts). (Inspired by 9router#1324.) (#5979)
  • feat(cli-tools): add CodeWhale — the actively-maintained successor to DeepSeek TUI (same author, renamed project) — as a dual dashboard entry alongside the existing "deepseek-tui" catalog entry, so existing DeepSeek TUI users keep a working card while new users are steered to CodeWhale. New /api/cli-tools/codewhale-settings route writes ~/.codewhale/config.toml and keeps the legacy ~/.deepseek/config.toml in sync. (Inspired by 9router#1761.) (#5996)
  • feat(server): support reverse-proxy basePath deployment via a new opt-in OMNIROUTE_BASE_PATH env var (empty by default), using Next.js's native basePath support so a deployment behind a reverse-proxy subpath (e.g. https://host/omniroute/) works without manual header stripping; the two hardcoded auth-redirect targets in src/server/authz/pipeline.ts now prefix with request.nextUrl.basePath. Default empty basePath is a no-op for existing root-path deployments. (Inspired by 9router#1810.) (#5992)
  • feat(providers): add SumoPod (ai.sumopod.com) and X5Lab (api.x5lab.dev) OpenAI-compatible BYOK aggregator gateways, wired via the default executor with bearer API-key auth; both use passthroughModels with a live /v1/models fetcher instead of a hardcoded catalog. Regression guard: tests/unit/sumopod-x5lab-provider.test.ts. (Inspired by 9router#1288.) (#5963)
  • feat(providers): add Charm Hyper (hyper.charm.land) as a new OpenAI-compatible, bearer-auth API-key gateway provider with a free tier (100 monthly Hypercredits); models resolve via passthrough (modelsUrl + live /v1/models) since the catalog isn't publicly documented. (Inspired by 9router#2006.) (#5961)
  • feat(providers): add Nube.sh (ai.nube.sh) as a new BYOK OpenAI-compatible gateway (LiteLLM proxy), Bearer/API-key auth. Its live model catalog is only reachable with a valid key, so no model IDs are hardcoded — it uses passthroughModels + modelsUrl for live enumeration. (Inspired by 9router#2294.) (#5936 — thanks @whale9820)
  • feat(providers): add b.ai (api.b.ai) as a new OpenAI-compatible BYOK provider, distinct from the existing thebai/theb.ai provider, using passthrough model discovery with no hardcoded model list. (Inspired by 9router#963.) (#5969)
  • feat(providers): add Qiniu (七牛云) AI inference gateway as a BYOK API-key provider — proxies many upstream models (DeepSeek V3/V4, Claude, Kimi, and more) behind a single key, shipping with an empty static seed and relying on passthroughModels + the live /v1/models catalog instead of a stale hardcoded model id. Regression guard: tests/unit/qiniu-provider.test.ts. (Inspired by 9router#911.) (#5966)
  • feat(providers): port ModelScope (Alibaba 魔搭) as a new API-key, OpenAI-compatible provider — verified against ModelScope's own docs that the real production domain is api-inference.modelscope.cn (.cn, not the upstream PR's .ai) and shipped passthroughModels: true with an empty seed + modelsUrl instead of the upstream PR's static 5-model snapshot, since the open-model catalog moves fast. (Ported from 9router#1764.) (#5965 — thanks @tn5052)
  • feat(providers): add Augment (Auggie CLI) as a new local, no-auth provider that spawns the user's local auggie CLI and pipes a flattened prompt via stdin, wrapping stdout as an OpenAI-compatible SSE stream or single JSON body. Auth is delegated to auggie login outside OmniRoute (synthetic noAuth: true connection, no DB row required); "Test Connection" spawns auggie --version. Hardened against the untrusted-input spawn sink: no shell: true on Windows (argv passed straight to the OS loader, no metacharacter interpretation), and model is validated against the registry allowlist before spawn (rejecting unknown or --prefixed values) with a trailing -- end-of-options marker. (Inspired by 9router#1200.) (#5972 — thanks @chamdanilukman)
  • feat(providers): add NVIDIA NIM image generation — a dedicated nvidia-nim image format/handler (separate host, ai.api.nvidia.com/v1/genai/<model>, native NIM body shape) for the 4 FLUX models (flux.1-dev, flux.1-schnell, flux.1-kontext-dev, flux.2-klein-4b), shaping each model's per-model request body (dimension/mode validation, required input image + aspect ratio, optional edit image) and normalizing the NIM response's varying shapes into the OpenAI {created, data} shape. (Inspired by 9router#1195.) (#5971)
  • feat(oauth): import a Codex connection from a raw ChatGPT access token — OmniRoute's only Codex import path previously required both access_token and refresh_token, leaving no path for a user with only a bare ChatGPT website access token. createProviderConnection gains an explicit access_token auth-type branch (intentionally never deduped), a new POST /api/oauth/codex/import-token route (Zod-validated), and OAuthModal's manual-paste path now detects an eyJ-prefixed pasted token and posts it to the new endpoint, mirroring the existing grok-cli raw-token flow. The executor's refreshCredentials() already degrades safely to null without a refresh token, forcing re-auth on expiry. (Inspired by 9router#1290.) (#5995 — thanks @ryanngit)
  • feat(dashboard): add a tool-source diagnostics settings toggle — a new Settings → Advanced card lets operators flip the existing logToolSources flag from the UI instead of editing the DB row directly; logToolSources is added to the .strict() /api/settings Zod PATCH schema (previously rejected). (Inspired by 9router#1825.) (#5978 — thanks @DuyPrX)
  • feat(dashboard): collapse and sort provider quota rows by remaining percentage — the expanded quota list is sorted highest-remaining-first and collapsed to the first 3 rows by default, with a "Show N more"/"Show less" toggle when a connection reports more than 3 quotas, keeping at-risk quotas visible above a long list of healthy ones. Sort/slice logic extracted into pure, directly-unit-tested helpers (sortQuotasByRemaining, getVisibleQuotas). (Inspired by 9router#1919.) (#5977)
  • feat(dashboard): suggest HuggingFace Hub media models — a new GET /api/v1/providers/suggested-models route proxies the public HF Hub models search API (Zod-validated, no token exposed client-side) and ImageExampleCard merges the results into the model picker as a selectable chip row for the huggingface provider; also adds a dedicated huggingface-image format/handler for HF's raw-image-bytes response. (Inspired by 9router#1633.) (#5990)
  • feat(cli-tools): add a Crush entry to the dashboard CLI-Tools catalog plus a new /api/cli-tools/crush-settings route (GET/POST/DELETE) — OmniRoute already shipped a crush CLI setup command (bin/cli/commands/setup-crush.mjs) but the dashboard catalog had no matching entry; the new route writes to the same canonical ~/.config/crush/crush.json path so the dashboard and CLI command agree. (Inspired by 9router#1233.) (#5970)
  • feat(providers): extend Vercel AI Gateway (vercel-ai-gateway/vag) beyond chat-only to support embeddings and image generation — the gateway's OpenAI-compatible /v1 API also exposes /embeddings and /images/generations, so entries were added to EMBEDDING_PROVIDERS (embeddingRegistry.ts) and IMAGE_PROVIDERS (imageRegistry.ts) modeled on the existing openai entries. (#5968 — thanks @tantai-newnol)
  • feat(api-keys): add per-key device/connection tracking — a SHA-256 fingerprint of IP + User-Agent, with a 30-minute TTL and per-key/global caps, tracks distinct client devices seen with each API key (in-memory only, raw IP never stored). A new GET /api/keys/[id]/devices route exposes masked device details, and the API Keys dashboard tab gets a "Devices" count badge alongside the existing Sessions badge. This is a new granularity distinct from the existing maxSessions cap, which limits concurrent sticky-routing sessions rather than tracking device identity. (#5998 — thanks @mugni-rukita)
  • feat(proxy): add Webshare (proxy.webshare.io) as a fourth source in the free-proxy provider framework alongside 1proxy, Proxifly, and IPLocate. WebshareProvider paginates the account's /api/v2/proxy/list/ endpoint, upserts proxies into the shared free_proxies table, and tombstones proxies the account no longer lists while never touching rows already promoted into the live proxy pool. Unlike the other sources, Webshare is a paid per-account list, gated on FREE_PROXY_WEBSHARE_API_KEY. (#5993 — thanks @ricatix)
  • feat(antigravity): support custom Google Cloud project ID settings from the connection edit modal (Antigravity family). (#5905 — thanks @nickwizard)
  • feat(dashboard): add a wildcard-CORS runtime warning banner (Settings → Authorization) when CORS_ALLOW_ALL/* origins are in effect, plus a new docs/security/CORS.md security guide covering the risk and safer alternatives. (#5602, #5759)
  • feat(api): add a /v1/audio/translations endpoint (Whisper-style audio translation), a new audioTranslation handler, and translation providers wired into audioRegistry. Regression guard: tests/unit/audio-translations-route.test.ts (8, incl. no-stack-leak). (#5809)
  • feat(providers): allow a custom icon URL for compatible provider nodes (migration 113 + nodes.ts + Zod schema + API routes + catalog + ProviderIcon UI). Regression guards: 14 backend + 5 frontend(vitest) + 24 page-utils tests. (#5815)
  • feat(xai): register a dedicated XaiExecutor with reasoning-effort suffix parsing. Regression guard: tests/unit/executors/xai-executor.test.ts (6). (#5800)
  • feat(webfetch): support self-hosted FireCrawl instances via FIRECRAWL_BASE_URL/FIRECRAWL_TIMEOUT_MS. Regression guard: tests/unit/executors/firecrawl-fetch.test.ts (4). (#5793)
  • feat(providers): add ClinePass as a first-class API-key (BYOK) provider — Cline's paid gateway (cline-pass/* models, plain Bearer key), distinct from the existing OAuth cline provider. Regression guard: 16 clinepass tests. (#5942 — thanks @adentdk)
  • feat(relay): gate Bifrost auto-routing by the provider plugin manifest — only manifest-eligible providers reach the sidecar; ineligible/unknown providers fall back to the existing TS routing path with explicit reasons. Regression guards: 4 provider-plugin-manifest + 11 relay-routing-backend tests. (#5870 — thanks @KooshaPari)
  • feat(providers): wire Claude Sonnet 5 end-to-end across the model pipeline — registries, modelSpecs, pricing (×3), cost, Sonnet-family fallback, 1M-context, and static models. (#5833 — thanks @ggiak)
🔧 Bug Fixes
  • dashboard (/dashboard/system/proxy 500 on every render): ProxyRegistryManager called useProxyBatchOperations(load) before the const load = useCallback(...) declaration in the component body, so every server render threw a TDZ ReferenceError: Cannot access 'load' before initialization and the whole proxy page 500'd (#5918 regression, caught by the release-PR e2e smoke — the PR→release fast-gates never render pages). The hook block now sits after the load declaration. Regression guard: tests/unit/ui/ProxyRegistryManager-tdz-render.test.tsx (SSR renderToString — the exact crash mode).

  • server (TRACE/TRACK/CONNECT returned raw 500 on every route): methods that undici/fetch cannot represent blew up inside Next's middleware adapter (TypeError: 'TRACE' HTTP method is unsupported.) as an unhandled 500 (caught by the release-PR dast-smoke Schemathesis negative tests on the new /api/keys/{id}/devices endpoint). The raw HTTP method guard now answers a clean 405 + Allow header for these methods on any path, before Next sees the request. Regression guard: tests/unit/dast-method-not-allowed.test.ts (new case).

  • i18n (auto-detect refreshed every first visit): LocaleAutoDetect (#5979) called router.refresh() on every cookie-less first visit — even when the detected browser locale was exactly the one the server had just rendered — re-navigating the page mid-interaction (flaky e2e "execution context destroyed" + a visible flash for every new visitor). It now refreshes only when the detected locale differs from the server-rendered <html lang>. Regression guard: tests/unit/ui/LocaleAutoDetect-refresh.test.tsx.

  • models (oc/ alias must reach the no-auth OpenCode provider): restore the #2901 routing contract after the #5918 transitive-alias change made the registered no-auth opencode provider unreachable by any prefix (oc/ chained through the manual opencodeopencode-zen slug override and misrouted its combo entries). resolveProviderAlias now stops the alias chain as soon as a hop lands on a registered provider id, while keeping #5918's transitivity across alias-only hops and its loop/depth guards. Regression guards: tests/unit/combo-builder-opencode-prefix.test.ts, tests/unit/provider-alias-transitive-5918.test.ts.

  • providers (Auggie executor EPIPE crash): a fast-exiting auggie CLI (e.g. binary present but immediately failing) delivered EPIPE asynchronously as an 'error' event on the child's stdin stream — which a plain try/catch around stdin.write() cannot catch — crashing the request instead of surfacing the sanitized CLI error. Both spawn sites now attach a stdin 'error' handler so the child's own exit/close handlers report the failure. Regression guard: tests/unit/auggie-executor.test.ts (deterministic 3/3 locally).

  • dashboard (CoolingConnectionsPanel broke next build): the cooling-connections panel from #6061 imported Card from a shadcn-style path that does not exist in this repo (@/components/ui/card) and pulled the server DB barrel (@/lib/localDb) into a client component — next build failed to compile on the release branch. The panel now renders with repo-native markup and reads formatResetCountdown from the new client-safe src/shared/utils/formatting.ts. Regression guards: tests/unit/format-reset-countdown.test.ts, tests/unit/ui/CoolingConnectionsPanel.test.tsx. (#6155)

  • oauth (Zed "Unknown provider" crash): adding Zed from the providers dashboard threw an unhandled OAuth GET error: Unknown provider: zed (500) (#6041). Zed is a keychain-import-only provider — it's listed in the OAuth catalog so the UI shows it, but has no OAuth handler, so the generic /api/oauth/[provider]/[action] route hit getProvider("zed") and crashed. The route now recognizes keychain-import-only providers and returns a clear 400 pointing users at the Import button (for both GET and POST OAuth actions), instead of a 500. Regression guard: tests/unit/oauth-keychain-import-only-6041.test.ts. (thanks @imblowsnow)

  • fix(providers): disable the unsupported thinking param for minimax-m2.7 on NVIDIA NIM (the upstream rejects it) (#6102). Regression guard: tests/unit/nvidia-minimax-thinking-strip.test.ts. (thanks @anki1kr)

  • fix(mitm): add an in-process guard so concurrent MITM server starts no longer race — a second start while one is already in flight is short-circuited instead of double-binding the listener (#6107). Regression guard: tests/unit/mitm-start-guard.test.ts. (thanks @anki1kr)

  • translator (Responses → Chat Completions): strip the Responses-API-only truncation field before forwarding a /v1/responses request to a non-OpenAI Chat Completions upstream (#6109). Strict upstreams (e.g. NVIDIA NIM) rejected it with HTTP 400 Unsupported parameter(s): truncation, breaking Codex-style clients routed to those providers. client_metadata, background, and safety_identifier were already stripped — truncation was the remaining gap. Regression guard: tests/unit/responses-strip-truncation-2311.test.ts. (thanks @TuanNguyen0708)

  • combo (prefer known context capacity over unknown): when a combo filters out at least one target for exceeding a known context limit, the router now prefers the remaining known-compatible targets over targets whose context metadata is simply unknown, instead of letting unknown-metadata targets be the only survivors. If no known-compatible context target remains, context-only candidates fall back to the normal strategy order. Regression guard: tests/unit/combo-context-window-filter.test.ts. (#6088 — thanks @Thinkscape)

  • models (GLM-5.2 context normalization): stop treating every hosted GLM-5.2 provider alias as the native 1M-context model. Native/bare GLM-5.2 and verified OpenCode / ZenMux routes keep their 1,000,000-token context, while hosted-provider aliases now respect the caps declared in their provider metadata instead of inheriting the native max. Regression guards: tests/unit/model-capabilities-registry.test.ts, tests/unit/models-catalog-route.test.ts. (#6091 — thanks @Thinkscape)

  • providers (Gemini Web): refresh the Gemini Web cookie handling and model catalog so live Gemini Web sessions keep authenticating and routing to current models. Regression guard: tests/unit/gemini-web.test.ts. (#6095 — thanks @backryun)

  • providers (Perplexity Web): refresh the Perplexity Web model catalog to the current set (GPT-5.4/5.5, Claude Sonnet 5.0 / Opus 4.8, GLM-5.2, Kimi K2.6, Nemotron 3 Ultra) and update the internal mode / model_preference mappings and thinking variants so requests resolve to live upstream models. Regression guard: tests/unit/perplexity-web.test.ts. (#6106 — thanks @backryun)

  • dashboard ("Update now" → Internal Server Error): clicking Update now on the dashboard home could crash the page with a blank "Internal Server Error" screen (Minified React error #31). The handler POSTs the loopback-only /api/system/version auto-update endpoint and, on a non-OK JSON response (e.g. a 403 when the dashboard is reached through a reverse proxy / non-loopback origin), passed the raw error envelope object { error: { code, message, correlation_id } } straight to notify.error(), which rendered the object as a React child and threw #31. The update-error path now funnels the body through extractApiErrorMessage() (the same safe extractor added in #5340), so a readable string always reaches the toast. Regression guard: tests/unit/ui/home-update-error-render-5991.test.ts. (#5991)

  • fix(onboarding): route the provider-details link in the onboarding wizard by the node's stable id instead of the composite provider slug, which could point at the wrong provider details page for multi-account/fingerprint nodes. Regression guard: tests/unit/onboarding-wizard-details-link-6145.test.ts. (#6145 — thanks @chirag127)

  • fix(cli): give setup-claude a fallback profile generator mirroring setup-codex, so profile generation no longer silently no-ops when the primary generator path is unavailable. Regression guard: tests/unit/cli/setup-claude.test.ts (new cases). (#6138 — thanks @derhornspieler)

  • fix(glm): suppress a leaked </think> close marker in the GLM Anthropic transport, which was surfacing the raw reasoning-close tag in visible response content instead of being consumed as part of the thinking-block framing. Regression guard: tests/unit/glm-think-close-marker-leak.test.ts. (#6133 — thanks @dhaern)

  • fix(provider-limits): close a TOCTOU race in quota-recovery clearing by moving the check-then-clear to a CAS (compare-and-swap) primitive in src/lib/db/providers.ts, so two concurrent recovery paths can no longer both observe stale state and double-clear/re-lock a connection. Regression guard: tests/unit/provider-limits-recovery.test.ts. (#6139 — thanks @janeza2)

  • fix(provider-limits): clear transient rate-limit state (rateLimitedUntil, lastError, backoffLevel) as soon as quota recovers, instead of leaving stale rate-limit fields behind that could keep a now-healthy connection looking unavailable. Regression guard: tests/unit/provider-limits-recovery.test.ts. (#6128 — thanks @janeza2)

  • combos (OpenCode/MiMo fingerprint accounts): expand fingerprint-scoped OpenCode/MiMo accounts into their full per-fingerprint set in the combo builder, which previously showed only the first matching account entry and hid the rest from combo target selection. Regression guard: tests/unit/combo-builder-fingerprint-expansion.test.ts. (#6092, closes #6087 — thanks @anki1kr)

  • fix(auth): persist quota-preflight account lockouts until the reset window elapses, instead of losing the lockout on process restart and letting a still-quota-exhausted account be selected again immediately. Regression guards: tests/unit/sse-auth.test.ts, tests/unit/opencode-quota-fetcher.test.ts, tests/unit/usage-service-hardening.test.ts. (#6090 — thanks @Thinkscape)

  • combo (fingerprint-based provider expansion): expand fingerprint-based providers into per-fingerprint combo targets (open-sse/services/combo/fingerprintExpansion.ts) so a combo referencing a fingerprint-scoped provider fans out to every matching fingerprint account instead of collapsing onto one. Regression guards: tests/unit/combo-fingerprint-expansion.test.ts, tests/integration/fingerprint-expansion.test.ts. (#6082 — thanks @pizzav-xyz)

  • fix (safety-net redirect reqId crash): fix a reqId ReferenceError thrown inside the safety-net combo redirect path in src/sse/handlers/chat.ts, remove dead code in src/domain/quotaCache.ts, and rename the stray root DESING.md to DESIGN.md. Regression guard: tests/unit/chat-safetynet-reqid-6097.test.ts. (#6097 — thanks @fix2015)

  • fix(compression): send a patch-only body to PUT /api/settings/compression from CompressionHub, instead of round-tripping the full settings object and risking clobbering fields changed elsewhere between load and save. Regression guard: tests/unit/ui/CompressionHub-patch-only.test.tsx. (#6077, closes #6039 — thanks @anki1kr)

  • fix(codex): use access_token.exp instead of id_token.exp when computing expiresAt on Codex auth import, since the id_token can expire far sooner than the actual access token, causing imported connections to be treated as expired while still usable. Regression guard: tests/unit/codex-auth-import-expiry.test.ts. (#6084, closes #6075 — thanks @anki1kr)

  • fix(security): persist the IP allow/block-list configuration (it was resetting to Disabled and clearing configured IPs on every restart/update) and actually enforce it in the authz pipeline (src/server/authz/pipeline.ts), where it was previously validated but never applied. Regression guards: tests/unit/ip-filter-persistence-6131.test.ts, tests/unit/authz/ip-filter-enforcement-6131.test.ts, tests/unit/ip-filter.test.ts. (closes #6131, #6132)

  • fix (Claude tool_result adjacency): reattach an OpenAI-shaped tool_result to sit directly adjacent to its originating tool_use before translating to Claude's message format (open-sse/translator/request/openai-to-claude/toolResultAdjacency.ts), since Claude's API rejects/mishandles a tool result separated from its tool call by intervening messages. Regression guard: tests/unit/translator-openai-to-claude.test.ts (new cases). (#6035 — thanks @KooshaPari)

  • fix(config): externalize ws/bufferutil/utf-8-validate in next.config.mjs so the copilot-m365-web executor's WebSocket masking path works at runtime — chat requests through it were silently timing out because the bundler was inlining ws instead of leaving it as a real Node dependency. Regression guard: tests/unit/next-config.test.ts. (#6130, closes #6062 — thanks @anki1kr, whose #6098 fix it re-lands)

  • fix(registry): update grok-cli model context lengths to match the actual Grok CLI /context capacities — grok-build 128k→256k, grok-composer-2.5-fast 128k→200k — so context-aware routing stops filtering these models out for exceeding a stale, too-low limit. Registry-only. (#5913 — thanks @Chewji9875)

  • fix(providers): strip an orphan tool_result (one with no preceding tool_use) on the Antigravity MITM path before translating to OpenAI format, since an unpaired tool result upstream caused request failures. Regression guard: tests/unit/antigravity-orphan-toolresult-6026.test.ts. (closes #6026, #6115)

  • fix(providers): emulate OpenAI-style tool_calls in the GitLab Duo executor (new open-sse/executors/gitlabResponses.ts), since the executor previously didn't emulate tool-call semantics for Duo, breaking tool-using clients routed to GitLab Duo. Regression guard: tests/unit/gitlab-duo-toolcalls-6051.test.ts. (closes #6051, #6111)

  • fix(429 / accountFallback): persist the per-account 429 cooldown cascade across the request boundary and classify OpenCode's "Monthly usage limit. Resets in N days." message as a connection-scoped quota exhaustion with an N-day cooldown (instead of a ~5s transient retry), so an exhausted account stops being re-selected until its window resets. (#6061 — thanks @KooshaPari / @anki1kr, whose superseded #6086 carried the same day-parser approach)

  • combo (sibling-model fallback on per-model-quota 500s): when a combo held multiple models from the same provider (e.g. two Gemini models) and the first returned a server 500, the router retried the same locked model and surfaced a 429 "cooling down" instead of trying the sibling — markConnectionLevelExhaustion was wrongly tripped by a model-level 500 for per-model-quota providers (gemini, github, passthrough, compatible), and the retry loop didn't check isModelLocked before re-hitting the same model. Both gaps are fixed; the combo now falls through to the untried sibling model. Regression guard: tests/unit/combo/combo-target-exhaustion.test.ts (21 cases). (#5976 — thanks @hartmark)

  • providers (Cline non-streaming envelope): Cline can return OpenAI-compatible chat completions wrapped as { success, data: { choices, usage, ... } }; the non-streaming path checked the top-level body for empty content before unwrapping, so a valid wrapped response could be misclassified as malformed/empty. The envelope is now unwrapped immediately after provider-envelope handling, before empty-content detection, usage extraction, and translation. Regression guard: tests/unit/cline-response-envelope.test.ts. (#6046 — thanks @KooshaPari)

  • providers (kimi-web, qwen-web): align the kimi-web model catalog and request-scenario selection with www.kimi.com's live GetAvailableModels response, and stop aliasing qwen3-coder-plus on qwen-web now that it is present as its own model in the live Qwen web catalog. (#5915 — thanks @janeza2)

  • translator (Antigravity/Gemini tool schemas): strip multipleOf from function-declaration parameters before forwarding to Antigravity/Gemini — it is not part of the Gemini OpenAPI 3.0 schema subset accepted upstream and triggered a hard 400 ("Unknown name multipleOf"). Added to GEMINI_UNSUPPORTED_SCHEMA_KEYS so it is stripped at every schema level; minimum/maximum are unaffected since Gemini accepts them. (Ported from 9router#2309, reported by @abil0321.) (#6052)

  • translator (Kiro system prompt leak): Kiro/CodeWhisperer has no system role, so system messages were normalized into a bare user turn — the full Claude Code system prompt then appeared as raw user text, polluting model context. System-origin content is now wrapped in <system-reminder> tags before merging into the Kiro user message; real user turns are unaffected. (Ported from 9router#2306, reported by @VitzS7.) (#6053)

  • fix(codex): convert Chat Completions json_schema response_format → Responses API text.format on the Codex path, and preserve an existing text.format through verbosity normalization. Regression guards: 48 translator-openai-responses-req + 8 codex-verbosity tests. (#5933 — thanks @yusufrahadika)

  • fix(thinking): only inject the redacted_thinking replay block when tool_use is present and thinking is enabled, avoiding a fabricated replay block on plain (non-tool) turns. (#5945, #5953)

  • fix(resilience): honor active codex session affinity over per-request reset-aware re-scoring, so an in-flight session sticks to its pinned account instead of being re-scored away mid-conversation. New src/sse/services/sessionAffinityPin.ts module. Regression guard: tests/unit/codex-session-affinity-reset-aware-5903.test.ts. (#5903, #5943)

  • fix(resilience): compute per-window is_exhausted and honor the quota-exhaustion preflight for priority combos, so a combo no longer keeps routing to a target whose current window is already exhausted. New open-sse/services/combo/quotaExhaustionCutoff.ts. Regression guard: tests/unit/combo-priority-quota-exhaustion-cutoff-5923.test.ts. (#5923, #5941)

  • fix(providers): strip a /v1 suffix from the base URL unconditionally in both models-discovery paths, avoiding a doubled /v1/v1/models fetch error (e.g. Api Airforce). Regression guard: tests/unit/airforce-v1-double-prefix-5899.test.ts. (#5899, #5920 — thanks @anki1kr)

  • fix(api): relax provider-scoped chat completion validation on /api/providers/[provider]/chat/completions. Regression guard: tests/unit/provider-scoped-chat-completions-validation.test.ts. (#5907 — thanks @nickwizard)

  • fix(providers): validate v0 Platform (Vercel) API keys via the /chats endpoint instead of a probe that rejected valid keys. Regression guard: tests/unit/provider-validation-specialty.test.ts. (#5954 — thanks @vittoroliveira-dev)

  • fix(mcp): auto-recover stale streamable HTTP MCP sessions on initialize instead of failing the reconnect. Regression guard: tests/unit/mcp-session-sweep.test.ts. (#5957 — thanks @Chewji9875)

  • fix(translator): enforce strict Anthropic content-block compliance when converting an antigravity → openai request. Regression guard: tests/unit/translator-antigravity-to-openai.test.ts (9). (#5935)

  • fix(sse): strip ANSI/VT100 escape codes from gemini-cli stream frames using a ReDoS-safe pattern. Regression guard: tests/unit/gemini-cli-ansi-sanitization.test.ts (5). (#5934 — thanks @anki1kr)

  • fix(discovery): resolve a doubled /v1 discovery path and a REDIRECT_BLOCKED probe-loop abort in the model-discovery route. Regression guard: tests/unit/provider-models-route.test.ts. (#5904 — thanks @hamsa0x7)

  • fix(providers): Perplexity Web now emits real tool_calls in streaming mode — previously only non-streaming requests (hasTools && !stream) converted <tool>{...}</tool> text into OpenAI tool_calls; streaming requests (the default for agentic coding clients) got the raw <tool> text as plain delta.content and never emitted a tool_calls SSE delta. Now mirrors the chatgpt-web toolMode helpers (buildToolModeResponse()/toolCompletionToSseStream(), extended with a caller-supplied idSeed so tool-call ids stay provider-specific), buffering the completion and emitting a terminal SSE replay carrying delta.tool_calls + finish_reason: tool_calls regardless of the caller's stream flag. (#5927, #5937)

  • providers (openai-family model inference no longer hijacks cataloged models): resolveModelByProviderInference() had an unconditional /^gpt-/i heuristic that hijacked any model id starting with gpt-/o1/o3 into provider openai, even when the id is cataloged under other providers — breaking bare (non-combo) requests for open-weight models like gpt-oss-120b (served by fireworks/cerebras/scaleway/byteplus/sambanova/heroku), which don't exist on openai's catalog, producing a 404 with no fallback. The heuristic is now gated on providers.length === 0 so it only fires for genuinely uncataloged openai-family ids. Regression guard: tests/unit/gptoss-provider-inference-5852.test.ts. (#5852, #5938)

  • fix(providers): deepseek-web reliability — auto-refresh the session on 401/403, refresh the v2.0.0 client headers, and fix the token-kind bulk import path. Regression guards: tests/unit/deepseek-web-autorefresh-401-response.test.ts, tests/unit/bulk-web-session-import.test.ts. (#5988 — thanks @backryun)

  • fix(api): guard the shared frontend API client (handleResponse in src/shared/utils/api.ts) against non-JSON error responses — it previously called response.json() unconditionally and read data.error directly, throwing an unrelated parse error (or undefined) instead of a useful message when an upstream/proxy returned a non-JSON error body. Now routes through parseResponseBody/getErrorMessage to build a safe message regardless of body shape. Regression guard: tests/unit/shared-api-utils.test.ts. (#5973)

  • fix(embeddings): forward the connection-level proxy configuration to embedding requests — src/lib/embeddings/service.ts previously ignored a connection's configured proxy when making embedding calls, so proxy-only network setups leaked embedding traffic outside the proxy. Regression guard: tests/unit/embeddings-proxy-forwarding.test.ts. (#5975)

  • fix(resilience): parse Retry-After from a 429's JSON body for cooldown calculation, not just the HTTP header — a new retryAfterJson.ts helper extracts a retry-after hint from common JSON error-body shapes and accountFallback.ts's cooldown path now prefers it when the header is absent. Regression guard: tests/unit/account-fallback-retry-after-json.test.ts. (Includes #6013's retry-after-json extraction.) (#5974 — thanks @KooshaPari)

📝 Maintenance
  • release close (release-PR one-pass CI sweep): restore Zod validation on the provider-scoped chat route with a .passthrough() schema that keeps #5907's relaxed semantics (t06 route-validation gate); point /api/keys/{id}/devices' 401 response at the management error envelope in docs/openapi.yaml (Schemathesis schema-conformance); rebaseline i18nUiCoverage.pct 77.5→76.8 (~1352 new en.json UI keys from the cycle await the async translation workflow — same shape as the v3.8.39 rebaseline); dismiss 2 CodeQL js/incomplete-url-substring-sanitization false positives on unit-test asserts (v3.8.35 precedent).

  • release close (Phase 0 pre-flight): align cycle-stale tests with merged behavior — provider count 166→167 (Kenari #6104), Linux-regenerated translate-path golden (+kenari), OpenCode quota scope providerconnection (#6061) — and absorb cycle ratchet drift (file-size caps for oauth/[provider]/[action]/route.ts 960, providerLimits.ts 998, chat.ts 1662, auth.ts 2426, with #6158 tracked to restore the oauth-route freeze). The test-masking gate gains a narrowly-scoped _deletedWithReplacement allowlist section (deletion is exempt ONLY when the declared replacement test file exists in HEAD — used for targetExhaustion.test.tstests/unit/combo/combo-target-exhaustion.test.ts, which has MORE coverage: 21 cases/52 asserts vs 13/37), plus 5 new gate unit tests and reduction-allowlist entries for the verified-legitimate #5958/#6088/#5816 assert migrations.

  • test (deflake setup-claude): tests/unit/cli/setup-claude.test.ts failed ~50% of runs with Unable to deserialize cloned data due to invalid or unsupported version at file teardown (all subtests passed), randomly reddening Unit Tests fast-path (2/2) / Fast Quality Gates across the PR→release queue. Root cause: node --test streams each file's report to the parent as V8-serialized frames on fd 1 (stdout), and the CLI helper under test (syncClaudeProfilesFromModels) prints progress via console.log — that stdout output interleaved with the serialized frames and corrupted the stream. The test now silences the stdout-writing console methods for the file's duration (no assertion inspects stdout), making it deterministic (15/15 green locally). (#5959) (#6021)

  • API validation: add a validatedJsonBody(request, schema) helper in src/shared/validation/helpers.ts that fuses JSON body parsing and Zod validation into a single call, returning either the type-narrowed data or a ready-to-return 400 NextResponse with the standard error envelope. Salvaged from the closed refactor PR #5075 (Tier 1 portable helper) with a focused 6-case regression test. Co-authored-by: KooshaPari KooshaPari@users.noreply.github.com

  • repo (Windows case-conflict cleanup): remove the stale root DESIGN.md, which case-conflicted with design.md and broke checkouts/clones on case-insensitive Windows filesystems. (#6140 — thanks @backryun)

  • i18n(zh-CN): translate the CHANGELOG entries and section headings, adopting zh-CN as a fully translated locale alongside the existing supporting docs. (#6043 — thanks @studyzy)

  • docs (env-doc-sync base-red): document BIFROST_PORT in .env.example / docs/reference/ENVIRONMENT.md — the Bifrost embedded-service merge referenced process.env.BIFROST_PORT (default 8080) without documenting it, so check:env-doc-sync failed on the release tip and reddened Fast Quality Gates for every open PR→release. Docs-only (8d7e3e28f).

  • test (CI-runner-independent translate-path golden): normalize OS/arch-derived request headers (X-Stainless-Os/X-Stainless-Arch, (OS;arch) User-Agent segments, and Antigravity's os.platform()-derived platform substring) in the provider translate-path golden snapshot, so the test no longer depends on the OS/arch of the CI runner that generated it — a Mac-literal Antigravity UA was failing on Linux CI. Regression guard: tests/unit/provider-translate-path-golden.test.ts. (#6076 — thanks @KooshaPari)

  • release-green base-reds (#5695 regex + file-size rebaseline): tests/unit/ui/quick-start-api-keys-link-5695.test.ts now tolerates Prettier splitting a multi-line <Link href=...> so the step1Desc regex matches the /dashboard/api-manager link instead of skipping to step2's single-line /dashboard/providers link (test was brittle, not the code). Also rebaselines 5 files that grew via already-merged release-tip PRs in config/quality/file-size-baseline.json (ApiManagerPageClient 3017→3058, OAuthModal 969→989, cliRuntime 1090→1100, webProvidersA 805→809, deepseek-web.test 1081→1092), with shrink tracked in #3501. (#6093)

  • release close (LEDGER-4 base-red): the cline-pass provider's minimax-m3 registry entry was missing supportsVision, breaking the LEDGER-4 registry-consistency test (every minimax-m3 entry must set supportsVision to match lite.ts — the model is multimodal). Flagged it to match every other minimax-m3 entry (trae, bazaarlink, cline, ollama-cloud, ...). (#6003)

  • release close (stryker tap.testFiles drift): additional release-green cleanup clearing the qoder registry's minimax-m3 supportsVision LEDGER-4 base-red and stryker.conf.json's tap.testFiles drift. (#6012)

  • install (pnpm 11+ support): pnpm 11 introduced ERR_PNPM_IGNORED_BUILDS for native addon packages — without explicit allowBuilds approval, packages silently skip their build scripts and OmniRoute fails to start with missing native modules. Sets allowBuilds=true for all 13 native addon packages in pnpm-workspace.yaml (@parcel/watcher, @swc/core, better-sqlite3, core-js, esbuild, keytar, koffi, libxmljs2, onnxruntime-node, protobufjs, sharp, tls-client-node, unrs-resolver) and migrates onlyBuiltDependencies from the deprecated package.json field to a new pnpm.json. (commit 39349da18 — thanks @chirag127)

  • refactor (Block J hot-path decomposition): extract pure leaves with no behavior change from the executor, translator, combo, and SSE hot paths — orphaned executor tests moved to top-level so a runner collects them, and handleComboChat's auto-strategy/target-timeout regions split into named helpers. (#6063, #6049, #6036, #6030, #6020, #6018, #6017, #6016, #6015, #6014, #6008, #6006, #6000, #5999, #5994, #5967, #5962, #5960, #5947, #5949, #5940, #5932)

  • chore (quality/CI housekeeping): rebaseline residual ESLint/cognitive-complexity/file-size drift accumulated over the v3.8.44 cycle, move orphaned executor tests to a top-level location so a runner actually collects them, harden the release pipeline with a test-masking pre-flight gate plus contributors/uncovered helpers, and make the pr-evidence FAIL output tell the author to push (a body edit alone does not re-run the gate). (#5926, #5944, #5952, #6027, #5928, plus a #5975-collateral test hardening pinning a seeded connection to direct egress in route-edge-coverage)

  • docs (housekeeping): normalize mixed-language documentation content, restore the OpenAPI coverage ratchet by documenting 9 newly-added routes, record Hard Rule #22 (cross-session safety — git stash + in-flight PR bans), and document the compression-engine's upstream sync policy for the RTK/Caveman engines. (#6105, #5955, #5948, plus docs-only commit 926b08aa8)

🙌 Contributors

Thanks to everyone whose work landed in v3.8.44:

ContributorPRs / Issues
@adentdk#5942
@anki1kr#5899, #5920, #5934, #6039, #6061, #6062, #6075, #6077, #6084, #6086, #6087, #6092, #6098, #6130
@artickc#6119
@backryun#5988, #6095, #6106, #6140
@chamdanilukman#5972
@Chewji9875#5913, #5957
@chirag127#6145
@derhornspieler#6138
@dhaern#6133
@doedjadirect commit / report
@DuyPrX#5978
@fix2015#6097
@ggiak#5833
@hamsa0x7#5904
@hartmark#5976
@imblowsnowdirect commit / report
@janeza2#5915, #6128, #6139
@KooshaPari#5870, #5974, #6035, #6046, #6050, #6061, #6073, #6076, #6086
@mugni-rukita#5998
@nickwizard#5905, #5907
@ofekbetzaleldirect commit / report
@pizzav-xyz#6082
@powellnormadirect commit / report
@ricatix#5993
@ryanngit#5995
@studyzy#6043
@tantai-newnol#5968
@Thinkscape#6088, #6090, #6091
@tn5052#5965
@TuanNguyen0708direct commit / report
@vittoroliveira-dev#5954
@waguriagenticdirect commit / report
@whale9820#5936
@WslzGmzsdirect commit / report
@yusufrahadika#5933
@diegosouzapwmaintainer
What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.43...v3.8.44

View originalPermalink
How v3.8.44 went

v3.8.43

Added 9
  • Usage endpoint and @@om-usage command now report personal API-key quotas as remaining percentages with provider USD cost drilldown via /api/usage/provider-window-costs endpoint
  • Usage quota system detects observed provider quota resets by same resetAt values instead of relying only on recorded weekly events
  • Live dashboard WebSocket can be fronted by reverse proxy or Cloudflare Tunnel via NEXT_PUBLIC_LIVE_WS_PUBLIC_URL environment variable with runtime support for prebuilt Docker and npm images
  • Handshake endpoint /api/v1/ws?handshake=1 now echoes lazily-read live.publicUrl for runtime WebSocket URL resolution
  • Optional auto-sync feature for CLI tool profiles after provider model sync, supporting Codex and Claude Code profiles via OMNIROUTE_AUTO_SYNC_CODEX_PROFILES and OMNIROUTE_AUTO_SYNC_CLAUDE_PROFILES feature flags
  • CLI Code dashboard now includes CLI profile auto-sync card to toggle Codex and Claude profile auto-synchronization
Changed 2
  • Provider quota remaining is scaled by configured quota cutoff so protected reserve reads as 0% left
  • Analytics surfaces now support opt-in flatRateAsZero cost option while budget, quota, and routing continue estimating unchanged for flat-rate providers
[3.8.43] — 2026-07-02
✨ New Features
  • usage (quota percentages + provider USD drilldown): @@om-usage and the HTTP usage endpoint now report personal API-key quotas as remaining percentages (USD amounts stay out of the command output), provider quota remaining is scaled by the configured quota cutoff so the protected reserve reads as 0% left, and the quota dashboard regains a provider USD cost drilldown (/api/usage/provider-window-costs + ProviderUsdCostModal, management-auth gated). Also honors observed provider quota resets: a same-resetAt reset (usage dropping back to the reset floor) is detected and preferred over stale recorded weekly events for provider USD windows and API-key USD quotas. New src/lib/usage/providerWindowCosts.ts. Regression guards: tests/unit/provider-window-costs.test.ts, tests/unit/internal-usage-command.test.ts, tests/unit/api-key-usage-limits.test.ts, tests/unit/lib/quota-reset-events.test.ts. Extracted from #5863 by @Witroch4.

  • dashboard (live WS behind reverse proxy): the live dashboard WebSocket can now be fronted by a reverse proxy or Cloudflare Tunnel via NEXT_PUBLIC_LIVE_WS_PUBLIC_URL (e.g. wss://ws.my-ai.com/live-ws). The URL is honored both at build time (env inlined into the bundle) and at runtime for prebuilt Docker/npm images: the /api/v1/ws?handshake=1 handshake now echoes a lazily-read live.publicUrl (only ws:///wss:// values are accepted; anything else is rejected to null), and useLiveDashboard resolves the URL from that handshake before connecting, falling back to the previous ws(s)://hostname:20129 default. Also documents LIVE_WS_ALLOWED_HOSTS and aligns the GitLab Duo OAuth scopes line in .env.example with the live config (ai_features read_user). Regression guard: tests/unit/live-ws-public-url.test.ts (5). (#5877 by @ianriizky)

  • providers (CLI profile auto-sync): opt-in toggles to auto-regenerate CLI tool profiles after a provider model sync. When enabled, a model-catalog change (re)writes that tool's profile files from the live catalog — Codex (~/.codex/*.config.toml) and now Claude Code (~/.claude/profiles/<name>/settings.json, via an extracted syncClaudeProfilesFromModels + a new claudeProfileAutoSync.ts mirroring the Codex path). Both are off by default and never touch the active/default CLI config; they are backed by the OMNIROUTE_AUTO_SYNC_CODEX_PROFILES / OMNIROUTE_AUTO_SYNC_CLAUDE_PROFILES feature flags (DB/dashboard override > env > default "false") and additionally gated behind the existing CLI_ALLOW_CONFIG_WRITES write-guard. A "CLI profile auto-sync" card on the CLI Code dashboard toggles each (moved from the providers dashboard in #5778 — thanks @rdself). Regression guards: tests/unit/claude-profile-auto-sync-gate.test.ts, tests/unit/codex-profile-auto-sync-gate.test.ts, tests/unit/cli/setup-claude.test.ts (follow-up to #5737).

  • cli (startup banner): the serve startup banner now prints the running OmniRoute version (v3.8.x) beneath the ASCII logo, so the active version is visible at a glance without a separate --version call. Regression guard: tests/unit/cli-serve-version-banner.test.ts. Thanks @chirag127 (#5752).

  • analytics (subscription cost): flat-rate providers now show $0 in cost analytics instead of an inflated per-token estimate. Subscription / coding-plan providers (every cookie-web provider — ChatGPT Web, grok-web, … — plus the dedicated Minimax Coding, Kimi Coding, GLM Coding, Alibaba Coding Plan, and Xiaomi MiMo plans) bill a flat fee, not per token, yet still carry per-token pricing rows used for estimates — so the analytics dashboard over-reported their cost. A new flat-rate classifier (src/lib/usage/flatRateProviders.ts) is consulted by the analytics surfaces (analytics route, usage stats, usage analytics) via an opt-in flatRateAsZero cost option, so those providers read $0 while budget / quota / routing keep estimating unchanged. Deliberately NOT zeroed: codex/cx (OmniRoute actively tracks Codex token cost — Fast-tier multipliers, GPT-5.x pricing — and Codex can be a metered account), byteplus (metered ModelArk), minimax-cn (metered China API). Regression guard: tests/unit/flat-rate-cost-5552.test.ts. (#5552)

  • mcp (RTK): expose the RTK tool-output learn/discover workflow as two new MCP tools so an agent can grow the RTK filter catalog without leaving the protocol. omniroute_rtk_discover analyzes recently captured raw tool output (discoverRepeatedNoise / suggestFilter) and returns candidate noise patterns plus a suggested filter; omniroute_rtk_learn lists the captured command samples (listRtkCommandSamples) and resolves a command to its RTK filter id (commandToId). Both are read-only (scope read:compression), wrap the existing RTK discovery primitives (no new logic in the engine), and log to the MCP audit trail. Regression guard: tests/unit/compression/rtk-mcp-tools.test.ts (4). gaps v3.8.42 — T07.

  • compression (LLM tier): add an opt-in, default-off LLM-tier compression engine (llm) that condenses the prose of non-system messages via a pluggable chat-completion backend. It mirrors the llmlingua engine's contract but is safe by construction: the default backend is a no-op pass-through (the engine never mutates the payload until an operator both enables it and wires a real backend via setLlmCompressorBackend()), it is not part of the default stacked pipeline, enabled defaults to false, fenced code blocks and system messages are never sent to the model, and every backend error fails open (the original segment/body is kept, never thrown). A minTokens floor skips small prompts. The real production backend is intentionally a VPS-validated follow-up (Hard Rule #18), exactly as the llmlingua worker backend is gated. New open-sse/services/compression/engines/llm/index.ts. Regression guard: tests/unit/compression/llm-compressor-engine.test.ts (8). gaps v3.8.42 — T05/C3.

  • memory (typed decay): add opt-in typed memory decay (TV6) so the conversational memory store stops accumulating stale episodic noise. Each injected memory now tracks an access_count + last_accessed_at (always-on, non-destructive telemetry; migration 111_memory_typed_decay), and an opt-in, default-off sweep (MEMORY_TYPED_DECAY_ENABLED, default false) deletes memories that are past a per-type TTL and not immune. Only episodic decays by default (30d, env-tunable); factual/procedural/semantic are immune, and any memory accessed >= 3 times earns access immunity (mirroring "guardrail/convention/decision never decay"). The decay clock re-bases on the last access, so used memories survive. Deletions reuse deleteMemory (SQLite + sqlite-vec + Qdrant stay in sync) and fail open; an optional periodic sweep is doubly opt-in (also needs MEMORY_TYPED_DECAY_SWEEP_INTERVAL>0). With the flag off nothing is ever deleted (Rule #20 spirit). New src/lib/memory/typedDecay.ts. Regression guard: tests/unit/memory/typed-decay.test.ts (15). gaps v3.8.42 — T10/TV6.

  • dashboard (combos): the named-combos editor now lets you drag to reorder the stacked-compression pipeline instead of only editing fixed-position steps. A new pure model (src/shared/components/compression/compressionPipelineModel.ts) owns add/remove/move/update with the engine→intensity invariant and a never-empty guarantee, and a @dnd-kit/sortable editor (CompressionPipelineEditor.tsx, matching the sidebar reorder pattern) replaces the inline list in CompressionCombosPageClient. Order persists through the existing combos endpoint. Regression guards: tests/unit/compression-pipeline-model.test.ts (11) + tests/unit/ui/compression-pipeline-editor.test.tsx (4). A dedicated tests/e2e/compression-studio.spec.ts (Tela A render + tab switch) closes the studios e2e gap the combo-live spec did not cover. gaps v3.8.42 — T06 + T03.

  • compression (pipeline): add an opt-in, default-off per-engine circuit-breaker to the stacked compression pipeline (T02). When an engine throws repeatedly across requests, its breaker opens and the stacked loops skip that engine (keeping the body verbatim for that step — fail-open) for a cooldown, then probe once (lazy half-open); success closes it, a failed probe re-opens it. This is distinct from the provider circuit-breaker (src/shared/utils/circuitBreaker.ts, provider-scoped + DB-persisted) — the new pipelineEngineBreaker.ts is engine-scoped, process-local, and adds zero DB/IO on the hot path. It composes with the existing per-request TV1 bail-out (which skips within a single request); the breaker adds cross-request memory. Default off (COMPRESSION_PIPELINE_BREAKER_ENABLED=false) → byte-identical to the pre-breaker pipeline (a throwing engine still propagates unless TV1 is separately enabled). Configurable per-call, per-CompressionConfig, or via env (_THRESHOLD/_COOLDOWN_MS). Regression guard: tests/unit/compression/pipeline-circuit-breaker.test.ts (9, incl. a throwing-engine integration); existing strategySelector/bail-out suites stay green. gaps v3.8.42 — T02 (2.2).

  • compression (CCR): the CCR retrieval-feedback (H8) is now graduated instead of a binary cliff. Previously a block retrieved >= 3 times was flagged do-not-compress and everything below that stayed fully compressible. Now each prior retrieval raises a block's effective minChars linearly (effectiveMinChars), so frequently-retrieved content is compressed progressively less; the >= 3 exclusion is preserved (as Infinity). The ramp is controlled by a retrievalRampFactor (default 2, per-combo config or COMPRESSION_CCR_RETRIEVAL_RAMP_FACTOR); 1 reproduces the exact legacy binary behavior. Per-(principal, hash) isolation is unchanged. Regression guard: tests/unit/compression/ccr-retrieval-ramp.test.ts (12); existing CCR suites (51) stay green. gaps v3.8.42 — T08/H8.

  • compression (cache-aware): add an opt-in, default-off usage-observed prefix freeze (H5). The cache-aware guard previously preserved the system prompt only for providers a static heuristic recognized as caching. It now also learns which system prompts actually recur: once a system prompt has been observed >= a threshold across requests, it is treated as a stable cacheable prefix and preserved from compression even for providers the static check misses — recovering prompt-cache hits that a prefix-compressing mode would otherwise bust. Content-addressed by a hash of the system prompt (OpenAI / Claude / Gemini shapes), in-memory + bounded, zero DB/IO; a "freeze" only preserves the prefix, so it never mutates a payload. Default OFF (COMPRESSION_PREFIX_FREEZE_ENABLED, threshold _THRESHOLD); respects the never preserve-mode (never freezes). New open-sse/services/compression/prefixFreeze.ts, wired into resolveCacheAwareConfig. Regression guard: tests/unit/compression/prefix-freeze.test.ts (10); 44 existing cache-aware / preserve-mode tests stay green. gaps v3.8.42 — T08/H5.

  • compression (read-lifecycle): add a new opt-in, default-off read-lifecycle engine (H7) that collapses stale/superseded file-Read tool results. In agentic conversations the same file is Read repeatedly; an earlier Read becomes stale once the same path is re-read (superseded by a newer view) or modified by a later Write/Edit. The engine replaces those earlier Read results with a short stub — keeping only the current (last, un-superseded) Read intact — recovering the tokens the model no longer needs. Unlike session-dedup (identical-content) or ccr (reversible markers), this is semantic + lossy, so it is opt-in (enabled defaults false). Conservative by construction: matches only well-known Read/Write tool names, compares exact paths, collapses a Read only when a strictly-later invocation touches the same path, and fail-opens on any unexpected shape. Supports both the Anthropic (tool_use/tool_result) and OpenAI (tool_calls + role:"tool") shapes. New open-sse/services/compression/engines/readLifecycle/index.ts. Regression guard: tests/unit/compression/read-lifecycle.test.ts (10). gaps v3.8.42 — T08/H7.

  • observability (correlation IDs): requests now carry a correlation id threaded through logs so a single request can be traced end-to-end across the pipeline. (#5834 — thanks @hartmark)

  • cli (startup banner — boot time): the serve ready banner now shows how long startup took, so slow-boot conditions are visible at a glance. (#5799 — thanks @ishatiwari21)

  • api (quota-policy bypass scope): add an opt-in API-key provider quota-policy bypass scope, so a designated key can be exempted from provider quota enforcement without disabling quotas globally. (#5731 — thanks @Witroch4)

  • providers (Ollama local): add a first-class Ollama local-provider card to the providers dashboard so the local LLM runtime can be configured like any other provider. (#5712 — thanks @diegosouzapw)

  • codex (fallback profiles): generate fallback CLI profiles for Codex-compatible models so compatible models get a usable profile automatically. (#5701 — thanks @skyzea1)

  • api (response-body validation + failover): add a configurable response-body validation step that can fail a target over to the next candidate when the upstream returns a structurally-invalid body (routing/#4985). (#5684 — thanks @diegosouzapw)

  • providers (SenseNova): complete the SenseNova free Token Plan — chat completions plus Text-to-Image (ported from 9router#2233). (#5679 — thanks @diegosouzapw)

  • db (self-correcting context windows): add self-correcting model context-window overrides so a model whose advertised context length is wrong is corrected automatically (models/#5004). (#5667 — thanks @diegosouzapw)

  • routing (latency strategy): optimize the latency routing strategy using observed per-target performance metrics for better candidate selection. (#5629 — thanks @KooshaPari)

  • compression (preserveSystemPrompt mode): add a preserveSystemPrompt mode enum (always | whenNoCache | never) with legacy back-compat, giving operators explicit control over when the system prompt is protected from compression (T05/C5). (#5653 — thanks @diegosouzapw)

  • commandCode (vision): add multimodal image support for Command Code vision models. (#5557 — thanks @Stazyu)

  • compression (read-lifecycle engine): T08/H7 (2.5) — an opt-in read-lifecycle engine that collapses superseded file reads so stale earlier reads of the same file are pruned from the context. (#5754 — thanks @diegosouzapw)

  • compression (usage-observed prefix freeze): T08/H5 (2.4) — opt-in prefix freeze driven by observed usage, keeping a stable cached prefix from being rewritten by downstream engines. (#5744 — thanks @diegosouzapw)

  • compression (CCR retrieval-feedback ramp): T08/H8 (2.3) — a graduated Context-Compression-Ratio retrieval-feedback ramp that tunes compression aggressiveness from retrieval signals. (#5739 — thanks @diegosouzapw)

  • compression (per-engine circuit breaker): T02 — an opt-in per-engine pipeline circuit-breaker that disables a misbehaving compression engine without failing the whole pipeline. (#5735 — thanks @diegosouzapw)

  • compression (LLM-tier engine): T05/C3 — an opt-in LLM-tier compression engine that uses a model pass for higher-ratio semantic compression. (#5702 — thanks @diegosouzapw)

  • dashboard (compression pipeline editor): T06/T03 — a drag-to-reorder compression pipeline editor plus a compression-studio e2e flow. (#5727 — thanks @diegosouzapw)

  • memory (typed decay): T10/TV6 — opt-in typed memory decay so aged, low-value memories fade on a per-type schedule. (#5723 — thanks @diegosouzapw)

  • mcp (RTK tools): T07 — expose the RTK learn/discover capabilities as first-class MCP tools. (#5691 — thanks @diegosouzapw)

  • providers (CLI profile auto-sync): opt-in CLI profile auto-sync toggles, including Claude Code auto-sync, so generated CLI profiles can track provider changes automatically. (#5755 — thanks @diegosouzapw)

🔧 Bug Fixes
  • fix(opencode): stop fabricating User-Agent: opencode/local and x-opencode-client: cli headers when the client sends none — the executor-dedup refactor (#5720) accidentally re-introduced header fabrication, violating the forward-only contract (inventing opencode-internal values risks upstream rejection). Restored to forward-only: those headers are emitted only when a real client source is present. Regression guard: tests/unit/opencode-executor.test.ts. (thanks @diegosouzapw)

  • fix(executors): resolveEffectiveKey returns undefined (not "") when no API key is present — a type-coercion cleanup (#5798) changed apiKey ?? "" to satisfy the typechecker, silently mutating auth-key resolution semantics. Widened the return type to string | undefined and reverted the coercion so OAuth-only credentials resolve correctly. Regression guard: tests/unit/refactor-buildHeaders-preamble.test.ts. (thanks @diegosouzapw)

  • fix(translator): restore the terminal message_delta + message_stop on Responses→Claude streams — the doubled-tool-args dedup (#5828) guarded the finish handler on the shared state.finishReason, which the openai-responses→openai leg sets first in the hub path, so the openai→claude leg dropped its terminal events and the stream ended after content_block_delta. The dedup now uses a dedicated state.claudeFinishEmitted flag. Regression guard: tests/unit/claude-code-rendering-fixes.test.ts. (thanks @diegosouzapw)

  • fix(pricing): add the Kiro claude-sonnet-5 pricing row so the newly-catalogued model (#5796) no longer reports $0.00 usage. Regression guard: tests/unit/catalog-updates-v3x.test.ts. (thanks @diegosouzapw)

  • fix(github): keep Copilot access-token sessions active. GitHub Copilot device-flow accounts can hold a GitHub access token plus a short-lived Copilot token without a refresh token; the proactive health check treated that as terminal no_refresh_token and marked the connection expired minutes after login. The health check now keeps those sessions active, clears stale no_refresh_token state, and refreshes the Copilot sub-token when needed. Regression guard: tests/unit/token-health-no-refresh-token-expired-5326.test.ts. Extracted from #5863 by @Witroch4.

  • fix(kiro): bound the Claude model-id dash→dot normalization to a 1–2 digit minor so date-suffixed ids (e.g. claude-opus-4-20250514) are no longer corrupted. (thanks @voravitl)

  • fix(usage): preserve (bounded) tool definitions in request logs even when the request body is truncated, so the request-details view can still show available tools. (thanks @noir017)

  • fix(providers): route OpenAI responses-only models to /v1/responses instead of 404ing on /v1/chat/completions. The curated gpt-5.5-pro / gpt-5.4-pro entries never worked (OpenAI only serves *-pro reasoning models via the Responses API), and "Test all models" surfaced the same 404s. The registry entries now carry targetFormat: "openai-responses" (reusing the existing per-model translation plumbing shared with gh/codex), DefaultExecutor.buildUrl swaps the openai endpoint to /responses in lockstep (honoring custom base URLs), and a -pro suffix heuristic covers dynamically-synced ids such as o1-pro / gpt-5.2-pro (same spirit as the gh executor's /codex/i routing, 9router#102). Legacy completions-only ids (e.g. gpt-3.5-turbo-instruct) are out of scope — they are not in the catalog and OmniRoute has no legacy /v1/completions upstream. Regression guard: tests/unit/openai-responses-only-models-5842.test.ts (8). Thanks @maikokan. (#5842)

  • fix(image): keep bare codex image aliases (e.g. gpt-5.5) resolving to the codex image pipeline even when a combo shares the same name. A chat combo named gpt-5.5 used to shadow the bare image alias in resolveImageRouteModel, hijacking /v1/images/* requests to a chat target (regression path adjacent to #5887); codex bare models are now reserved before bare-combo resolution, while non-codex aliases (e.g. gpt-image-2) remain user-shadowable (#3214/#3215 behavior preserved). Regression guard: tests/unit/image-routes-combo-edits-3214-3215.test.ts (9). (#5902 by @KooshaPari)

  • fix(ci): re-green the release/v3.8.43 fast-gates queue — every PR→release was inheriting base-reds (#5798). Five distinct blockers cleared: (1) stale modelContextOverrides entry in the check:db-rules intentionally-internal allowlist (#5827 allowlisted it while the #5609 fix re-exported it from localDb.ts; the re-export stays, the obsolete entry goes, classification guard re-pinned to 33); (2) LIVE_WS_ALLOWED_HOSTS / NEXT_PUBLIC_LIVE_WS_PUBLIC_URL documented in docs/reference/ENVIRONMENT.md (env/docs contract, from #5877); (3) the Router Backends ADR's references to the not-yet-merged registry (#5868) marked as landing-with-PR so check:fabricated-docs --strict passes; (4) antigravity-429-quota-tdd + middleware-header-strip-5849 added to stryker tap.testFiles (check:mutation-test-coverage); (5) file-size / complexity / cognitive-complexity ratchets rebaselined with justification notes — all drift measured identical on the pristine tip and this PR (net-zero). Regression guard: tests/unit/check-db-rules-classification.test.ts. (#5798)

  • providers (codex image auto-routing regression): an unprefixed gpt-5.5 request from a codex-only setup (no OpenAI connection) now correctly infers the codex provider again — the OpenAI static-catalog short-circuit in resolveModelByProviderInference was preempting the codex-preference block, so gpt-5.5 (added to the OpenAI catalog) stopped auto-routing to Codex image generation. Users with an active OpenAI connection are unaffected (OpenAI stays default). Regression guard: tests/unit/codex-gpt55-routing-5887.test.ts. (#5887)

  • api (proxy header hygiene): upstream x-middleware-* control headers (emitted by providers hosted behind Next.js, e.g. synthetic.new) are now stripped from proxied responses instead of forwarded verbatim — forwarding x-middleware-rewrite made Next 16 throw NextResponse.rewrite() was used in a app route handler and return 500 despite a successful upstream call. Applies to both streaming and JSON paths. Regression guard: tests/unit/middleware-header-strip-5849.test.ts. (#5849)

  • docs (pnpm global install): replaced the unsupported pnpm approve-builds -g step with the install-time pnpm add -g omniroute@latest --allow-build=better-sqlite3 flag across README + Setup Guide (and i18n mirrors), fixing native-build approval for pnpm v11 global installs. (#5554)

  • dashboard (token badge): the red "Token Expired" connection badge no longer flashes for OAuth refresh-capable providers (Antigravity/Gemini) whose access token merely lapsed but is auto-refreshed — it now shows only when the connection is terminally expired (testStatus === "expired"). Continuation of #5326. Regression guard: tests/unit/ui/connection-row-token-badge-5836.test.tsx. (#5836)

  • db (auto backup toggle): full pre-write SQLite backups now honor the persisted backup.autoBackupEnabled dashboard setting — previously only the DISABLE_SQLITE_AUTO_BACKUP env var was checked, so disabling auto-backup in the UI had no effect and ~70MB pre-write snapshots kept firing. Manual and pre-restore backups still always run. Regression guard: tests/unit/db-backup-autobackup-setting-5871.test.ts. (#5871)

  • providers (auto/ routing for custom providers): custom OpenAI-/Anthropic-compatible providers (dynamic *-compatible-* connection IDs) are no longer excluded from auto/ routing — the Auto-Combo virtual factory previously skipped any connection whose provider was absent from the static registry. It now falls back to the connection's defaultModel. Regression guard: tests/unit/auto-custom-provider-5873.test.ts. (#5873)

  • middleware (hook sandbox): operator-authored pre-request hook code now runs inside a hardened Node vm sandbox (minimal context, no ambient globals/process.env, execution timeout, no require) instead of new Function() in the main process — closing the Hard Rule #3 / SonarCloud S1523 exposure. Regression guard: tests/unit/middleware-hook-sandbox-5872.test.ts. (#5872)

  • mcp-server (auth forwarding): the per-caller MCP identity forwarded via withMcpHttpAuthContext now wins over the static OMNIROUTE_API_KEY env fallback in the internal-fetch helpers (apiFetch, omniRouteFetch) — previously the env key was spread after the forwarded headers and clobbered the caller's Authorization. Regression guard: open-sse/mcp-server/__tests__/httpAuthContext.test.ts. (#5819)

  • dashboard (Modal provider — two-field auth): the Modal provider connection form now exposes two fields — Token ID + Token Secret — instead of a single API-key input, since Modal authenticates with Authorization: Bearer <token-id>:<token-secret>. The dashboard combines the two fields into the id:secret credential before saving (combineModalCredential, trims both parts), while a value pasted in the legacy single-field format keeps working verbatim (empty secret → passthrough), so existing saved connections need no migration; the key-help link points at Modal's token settings. Regression guard: tests/unit/modal-credential-combine.test.ts (5). (#5881, closes #5446) Follow-up: the Validation Model Id field is now pre-filled for Modal with the same model the server-side validator probes (Qwen/Qwen3-4B-Thinking-2507-FP8, shared via MODAL_DEFAULT_VALIDATION_MODEL_ID in src/shared/constants/modal.ts), closing the last checklist item of #5446. Regression guard: tests/unit/modal-validation-model-prefill.test.ts.

  • api (chat completions — early SSE keepalive gate): the /v1/chat/completions route wrapped the response in the early-stream keepalive whenever stream was not explicitly false, so a client that omitted stream and asked for JSON (Accept: application/json) could receive premature SSE framing. The keepalive wrapper is now gated on an explicit stream: true in the body or an Accept header that forces SSE (acceptHeaderForcesStream); the parsed body is passed to the chat handler untouched, so the actual stream/JSON framing stays decided by chatCore/resolveStreamFlag — preserving OmniRoute's legacy streaming default when stream is omitted and the per-key streamDefaultMode: "json" opt-in. Regression guard: tests/unit/chat-combo-live-test.test.ts ("returns JSON without early SSE framing when stream is omitted and Accept is application/json"). (#5866 by @rdself)

  • fix(github): drop a trailing assistant prefill before dispatching to GitHub Copilot chat to avoid 400 errors. (thanks @baslr)

  • fix(oauth): prevent cross-IdP account overwrites by disambiguating OAuth connections on username when present, not email alone. (thanks @KunN-21)

  • fix(mitm): best-effort revert privileged /etc/hosts entries on exit when a sudo password is cached, instead of always leaving orphaned state. (thanks @manhdzzz)

  • providers (Kiro — Claude Sonnet 5): the Kiro provider's model catalog was missing claude-sonnet-5, so the model could not be selected or routed even on accounts that already had access to it ("claude-sonnet-5 is not supported"). Added the model to the Kiro registry (open-sse/config/providers/registry/kiro/index.ts) as a 1M-context / 128K-output Claude model, mirroring the existing Claude entries; the registry models[] feeds both the model selector and the live CodeWhisperer ListAvailableModels fallback, so the model is now selectable and routable. Regression guard: tests/unit/kiro-claude-sonnet-5-2267.test.ts. (thanks @openbioinfo)

  • settings (model aliases — self-heal after restart): the Settings → Routing page showed "No exact-match aliases configured" after a server restart even though the aliases were persisted in the DB. Aliases are held in a module-local _customAliases map in modelDeprecation.ts that the boot path hydrates, but Next.js compiles the app-route module graph separately from the startup graph (the same webpack chunk-splitting class as #5312), so the GET /api/settings/model-aliases handler read a different, un-hydrated copy. The handler now self-heals: when its in-memory alias map is empty it reads settings.modelAliases from the DB (via the existing getSettings() db module — no raw SQL in the route) and repopulates the map, so the UI reflects the persisted aliases on the first GET after a restart. Follow-up: the root cause is now also fixed — the _customAliases store in modelDeprecation.ts is backed by globalThis (key __omniroute_customAliases__), so the startup and app-route module graphs share one store and the route reads the boot-hydrated aliases directly (the DB self-heal remains as a harmless fallback), mirroring the same globalThis singleton pattern already applied to thinkingBudget.ts/backgroundTaskDetector.ts (#5312). Regression guards: tests/unit/model-aliases-settings-route-selfheal.test.ts + tests/unit/model-aliases-globalthis-5777.test.ts. (#5777 — thanks @jleonar2)

  • providers (grok-cli token auto-refresh): grok-cli OAuth tokens were never proactively refreshed before their real expiry. mapTokens hardcoded expiresIn: 21600 (6 h) regardless of the token's actual lifetime, so the persisted expiresAt was always "now + 6 h" and the proactive tokenHealthCheck sweep (refresh when expiresAt - now < 5 min) fired 6 h after import instead of shortly before the token really expired. mapTokens now computes expiresIn from the authoritative expires_at field in ~/.grok/auth.json (ISO → epoch-seconds) with a fallback to the JWT exp claim (payload-only decode, no signature trust); the hardcoded 21600 is kept only when neither is present. An already-expired token (real expires_at/exp in the past) is now clamped to a positive expiresIn via Math.max(1, …), so the import route stores a near-future expiresAt and AutoCombo refreshes the connection instead of reading a past date and excluding it outright. Regression guards: 5 cases in tests/unit/grok-cli-oauth.test.ts (JWT exp, JSON expires_at, the 21600 fallback, and the two expired-token clamps). (#5775 — thanks @Chewji9875)

  • compression (CCR retrieve via MCP HTTP): the omniroute_ccr_retrieve MCP tool returned "CCR block not found" for blocks stored earlier in the same session when called over the MCP HTTP transports (SSE / Streamable HTTP), e.g. from OpenCode in a Docker deployment. Compression stores each block keyed by the API-key principal (String(apiKeyInfo.id)), but the tool resolved the caller via extra.authInfo.clientId — which the MCP SDK never populates for API-key auth — so it fell back to "anonymous" and the compound store-key never matched. The retrieve tool now resolves the caller's API-key id from the MCP HTTP auth context (httpAuthContext) using the same getApiKeyMetadata lookup used at storage time, so retrieval matches storage. Cross-tenant IDOR isolation is preserved: a different key resolves to a different id → miss; no key → the anonymous bucket only. Regression guard: tests/unit/compression/ccr-mcp-principal-5649.test.ts (extraction, distinct-principal isolation, fail-closed, end-to-end store→retrieve). (#5649)

  • compression (context-editing telemetry): streaming responses now record Context Editing savings. Anthropic surfaces context_management.applied_edits[] on the final message_delta snapshot of an SSE stream, but the streaming reconstruction (buildStreamSummaryFromEvents → Claude branch) dropped context_management entirely and no telemetry hook was wired into the streaming finalizer — so the delegated server-side context-clear savings (cleared_input_tokens / cleared_tool_uses) surfaced under engine context-editing in compression analytics only for non-streaming responses. The collector now preserves context_management from the final snapshot (last-writer-wins), and onStreamComplete mirrors the non-streaming recordContextEditingTelemetryHook (best-effort, Claude-only, HTTP 200 only). Purely additive telemetry — no payload mutation, no new env flag, no behavior change when the stream carries no context_management. Regression guard: tests/unit/context-editing-streaming-telemetry.test.ts (3). gaps v3.8.42 — T01 (5.1).

  • proxy (relay test diagnostics): the Proxy Pool "Test" button showed a bare "failed" with nothing in the server logs when a relay (Vercel / Deno / Cloudflare) responded with a non-200 — e.g. a 401 from an auth-token mismatch after a STORAGE_ENCRYPTION_KEY rotation. The relay success-path response set success: false but carried no error field, so the dashboard had no reason to show and the server logged nothing. The test now returns an actionable error (the HTTP status, plus an auth/encryption-key hint on 401/403) and logs the failure server-side; the SOCKS5/HTTP proxy path now logs its failures too. Shaping extracted to buildRelayTestResult with a regression guard (tests/unit/proxy-relay-test-error-5716.test.ts). Note: this surfaces why a relay fails — it does not repair a genuinely broken/misconfigured relay. (#5716)

  • fix(dashboard): add error boundaries for the Combos and MITM Proxy pages so a render error shows a recoverable fallback instead of a blank page. (thanks @wahyuzero)

  • providers (onboarding wizard — unsupported validation): adding a provider whose credentials have no live validator (LMArena, PiAPI, …) failed silently in the Add-Provider wizard. The /api/providers/validate endpoint returns HTTP 400 + { unsupported: true } for these (#5565/#5567), but the wizard's validateOnboardingApiKey ran it through expectOk, which threw on the non-200 — so the flow jumped to the error step and the connection was never created. The wizard now treats unsupported: true as a non-blocking "can't verify" and proceeds to save, mirroring AddApiKeyModal. Regression guard added to tests/unit/provider-onboarding-wizard.test.ts. (related to #5692)

  • dashboard (Quick Start step 1): the Quick Start "Create API key" step told users to "Go to Endpoint → Registered Keys" and linked to /dashboard/endpoint, but API keys are created on the API Manager page (/dashboard/api-manager, sidebar "API Keys") — the Endpoint page has no "Registered Keys" section, so users followed the link and could not find where to create a key. Step 1 now reads "Go to API Keys" and links to /dashboard/api-manager. Regression guard: tests/unit/ui/quick-start-api-keys-link-5695.test.ts. (#5695)

  • providers (DashScope/Alibaba setup link): the "Get API key" link for the Alibaba and Alibaba (China) providers pointed at the bare API host (dashscope-intl.aliyuncs.com / dashscope.aliyuncs.com), which returns 404 in a browser — API hostnames have no homepage. Repointed to the consoles where keys are actually issued: bailian.console.alibabacloud.com (international) and dashscope.console.aliyun.com (China). Same class as #5572/#5574/#5576; regression guard added to tests/unit/provider-setup-links-5572.test.ts. (#5665)

  • thinking / runtime-config (module-graph fix): operator-configured proxy settings that are hydrated at boot but read per-request were silently ignored in production. Next.js compiles instrumentation.ts (boot hydration via applyRuntimeSettings / restore hooks) as a separate webpack module graph from the app-route / open-sse executors, so a module-local let _config singleton is duplicated — the boot copy is hydrated but the request path reads a different, un-hydrated copy. Live VPS validation proved the Thinking-Budget hydration ran to completion at boot yet base.ts still saw the passthrough default (this is why #5312 fix A stayed broken even after the boot-wiring fix). Fixed by backing the singletons with globalThis (the pattern systemPrompt.ts already uses for the Global System Prompt, #2470), so all module-graph copies share one instance: thinkingBudget.ts (the dashboard Thinking-Budget mode now reaches the executor), backgroundTaskDetector.ts (the opt-in background-model degradation now actually fires on requests), and systemTransforms.ts (operator pipeline overrides now reach the request path). payloadRules.ts was already safe (it lazily self-loads from the DB per request, #2986). Regression guards: tests/unit/thinking-budget-globalthis-5312.test.ts + tests/unit/runtime-config-globalthis-5312.test.ts (assert globalThis-backed sharing; a module-local let fails them). (#5312)

  • thinking (Claude OAuth): restore the proxy-level Thinking-Budget config on startup. The dashboard mode (auto/custom/adaptive) is persisted under settings.thinkingBudget, but the boot-time hydration (hydrateThinkingBudgetConfig) was only wired into src/server-init.ts — an unused module that never runs in production — so the operator's choice silently reverted to the passthrough default on every restart (#5312 fix A was non-functional, even though its direct unit test passed). The hydration now runs in the real boot path (src/instrumentation-node.ts), alongside the Global System Prompt restore. Surfaced by live Anthropic-OAuth validation on the VPS. Regression guard: tests/unit/thinking-budget-boot-wiring-5312.test.ts (asserts the production boot module calls the hydration, not just the function in isolation). (#5312)

  • translator/chatcore (hardening): re-apply two defensive review-fixes that were dropped in a branch rebuild before #5661 / #5662 landed. (1) mergeConsecutiveSameRoleContents (OpenAI→Gemini) now shallow-copies each entry and its parts array instead of pushing the input reference, so the consecutive-same-role merge never mutates the caller's objects. (2) defaultClaudeToolType (Claude tool defaults) now passes any non-object array entry (null / primitive) through unchanged instead of spreading it into a fabricated { type: "custom", … } tool. No behavior change on real payloads (Gemini contents are freshly built; Claude tools are always objects); both properties are now locked by regression tests in tests/unit/translator-gemini-consecutive-role-2191.test.ts and tests/unit/claude-tool-type-default-2195.test.ts.

  • providers (grok-cli): truncate the tool list when it exceeds a provider's hard limit, so grok-cli (cli-chat-proxy.grok.com, max 200 tools) no longer rejects requests with Maximum tools limit reached. Adds a proactive PROVIDER_TOOL_LIMITS map (grok-cli: 200, consulted before the reactive cache), a corrected limit-parsing regex that captures the stated maximum (200) instead of the supplied count (427), and removes the broken < MAX_TOOLS_LIMIT truncation gate so truncation now fires whenever tools.length exceeds the effective limit. Regression guard: tests/unit/tool-limit-detector.test.ts. (#5563 — thanks @Chewji9875)

  • resilience (antigravity): record model lockout for Antigravity 429 rate_limit_exceeded errors. Antigravity's "Resource has been exhausted (e.g. check quota)." text was matched by overly broad QUOTA_PATTERNS and misclassified as QUOTA_EXHAUSTED, so the combo retry path was skipped (providerExhausted) and the model was never cooled down. Classification now prefers the structured error code — classifyErrorText(structuredError?.code || errorText) — so a rate_limit_exceeded code is treated as a transient rate-limit (not quota), and the two broad patterns (/resource.*exhaust/i, /check.*quota/i) were replaced with Antigravity-specific ones (individual quota reached, enable overages). (#5579 — thanks @Chewji9875)

  • providers (OpenAI-compatible): Codex MCP / tool_search deferred discovery (and apply_patch) now works through a Custom OpenAI-compatible provider. When such a provider received a Responses-API-shaped request that carried MCP / tool_search tools, OmniRoute downgraded it to /chat/completions, which drops the deferred tool-discovery mechanism — so the MCP namespaces never surfaced to the model and apply_patch was mis-handled as a JSON tool. The executor now detects a Responses-shaped request (input / previous_response_id / max_output_tokens / reasoning) that carries namespace / tool_search* tools and routes it to the upstream /responses endpoint natively instead of downgrading (it can also be forced via providerSpecificData._omnirouteForceResponsesUpstream). This is a distinct code path from the official Codex OAuth backend (#3033 / #4539, which the earlier fix never touched). Regression guard: tests/unit/executor-default-base.test.ts. Thanks to @KooshaPari for the fix. (#5483)

  • dashboard (routing): selecting the fusion strategy on the Global Routing defaults tab now reveals fusion-specific config instead of only the generic resilience fields. Fusion's engine knobs — judgeModel (the model that synthesizes the panel answers) and fusionTuning (minPanel / stragglerGraceMs / panelHardTimeoutMs) — already existed in the schema and the per-combo editor, but the Global Routing tab never surfaced them, so picking "fusion" there was effectively a no-op. The fields are now shown (extracted into a new FusionDefaultsFields component). Voting / aggregation-mode / per-provider-weight are intentionally not shown — those don't exist in the fusion engine. Regression guard: tests/unit/ui/combo-defaults-fusion-5598.test.tsx. (#5598)

  • dashboard (free proxy pool): the free proxy pool "Sync All" no longer fails silently with Total: 0. Three fixes: (1) the IPLocate source fetched …/protocols/<proto>.json and parsed it as JSON, but the upstream list is plain text (<proto>.txt, one ip:port per line) — every protocol 404'd / failed to parse; it now fetches .txt and parses the line list. (2) The sync route isolates each source in its own try/catch, so one provider throwing (e.g. a TLS handshake failure) no longer aborts the whole sync — the working sources still populate the pool. (3) The UI now surfaces the per-source errors the route already returns, instead of discarding the response, so a partial/empty sync explains itself. Regression guards: tests/unit/free-proxy-providers.test.ts, tests/unit/proxy-pool-sync-4878.test.ts, tests/unit/free-pool-tab.test.tsx. (#5595)

  • dashboard (memory engine): the memory engine status page no longer mixes English and Portuguese. The embedding / vector-store / rerank status detail strings were hardcoded in Portuguese in the backend (resolveEmbeddingSource, engineStatus), e.g. auto: nenhuma fonte de embedding disponível and sqlite-vec ativo, dim=…, while the surrounding UI labels render from the English i18n bundle — so an English user saw a half-translated page. The backend detail strings are now English (auto: no embedding source available, sqlite-vec active, dim=…, etc.), matching the rest of the page. Regression guard: tests/unit/memory-engine-status.test.ts. (#5596)

  • providers (cline): stop falsely mapping valid Cline (OAuth) responses to 502 empty_choices + account cooldown. detectMalformedNonStream only recognized choices[].message.content as a string, but some OpenAI-compatible upstreams — Cline via OAuth among them — return content as an array of Anthropic-style text blocks inside an OpenAI envelope. A non-empty response (recvBytes > 0) was therefore classified as empty_choices and turned into a 502 that also cooled the account down. The malformed-response detector now also treats a content array carrying at least one non-empty text block as real output. Regression guard: tests/unit/diagnostics.test.ts. (#5559)

  • embedded services (Windows): fix CLIProxyAPI install failing instantly with spawn unzip ENOENT on Windows. The binary extractor spawned unzip, which is not a Windows system command — it only ships inside Git for Windows' usr/bin, a directory Node's spawn PATH never sees, so even users with Git installed hit the error. On Windows the extractor now uses PowerShell's built-in Expand-Archive (via execFileAsync, no shell — paths pass as a single non-interpreted arg, with ''-escaping + -LiteralPath as defense in depth); other platforms keep using unzip. This is distinct from #5379 (that was npm.cmd needing shell: true). Regression guard: tests/unit/binary-manager-extract-zip-5590.test.ts. (#5590)

  • storage (daemon): fix a Node.js out-of-memory crash on startup when storage.sqlite grows large (~170 MB+). The boot-time call-log cleanup (cleanupExpiredLogsrotateCallLogs) ran two unbounded SELECT … FROM call_logs … .all() queries — listReferencedArtifacts (every artifact path) and deleteCallLogsBefore (every id before the retention cutoff). node:sqlite's StatementSync.all() materializes the entire result set as JS objects at once, so on a large table the V8 heap blew up and the process crashed before binding (FATAL ERROR: … heap out of memory, native frame node::sqlite::StatementSync::All). Both queries now page through call_logs in bounded 5 000-row chunks (new src/lib/usage/callLogsBoundedQueries.ts), keeping peak memory flat regardless of table size — no more manual --max-old-space-size bump required. Regression guard: tests/unit/call-log-oom-unbounded-5618.test.ts. (#5618)

  • dashboard (provider setup): fix three provider setup links that pointed at 404 pages. Ollama Cloud / ollama-search linked to ollama.com/settings/api-keys → corrected to ollama.com/settings/keys (the page moved; Ollama Cloud is a real keyed service, so the field stays). SearchAPI linked to the bare searchapi.io/docs (404) → searchapi.io/docs/google. You.com linked to you.com/docs/search/overview (404) → you.com/business/api/ (the developer portal). All three replacements were verified live. Regression guard: tests/unit/provider-setup-links-5572.test.ts. (#5572, #5574, #5576)

  • providers (AI/ML API): the model-import step now loads the live AI/ML API catalog (400+ models) instead of falling back to a stale 6-model seed. The registry had no modelsUrl, so the route silently used the bundled catalog with an "API unavailable — using local catalog" warning even when the key was valid. AI/ML API exposes its full catalog at the public, auth-free https://api.aimlapi.com/models endpoint (a bare array of { id, type, info }, distinct from the OpenAI-compat /v1/models); it's now wired into the models route's discovery config, with the bundled catalog kept as the offline fallback. Regression guard: tests/unit/provider-models-route.test.ts. (#5570)

  • providers (CablyAI): mark CablyAI deprecated — cablyai.com no longer resolves (DNS NXDOMAIN, verified 2026-06-30); the domain is gone. The provider is removed from the models-route discovery config so the import step returns a clean error instead of an unhandled 500 crash (the dead-domain fetch threw with no local-catalog fallback), and the registry entry now carries deprecated: true / riskNoticeVariant: "deprecated" so the dashboard flags existing connections (same treatment as the shut-down glhf/kluster.ai gateways). Regression guard: tests/unit/provider-models-route.test.ts. (#5568)

  • dashboard (provider add): non-LLM search/agent providers no longer fail the model-import step with a red Provider <id> does not support models listing. Jules (Google Labs coding agent), linkup-search (Linkup web search), ollama-search (Ollama Cloud web search — distinct from the local Ollama LLM), and searchapi-search (SearchAPI SERP) have no /v1/models endpoint, so the import surfaced a failure for expected behavior. Each now ships a small static catalog of its selectable capability ids — Linkup's fast/standard/deep search depths, SearchAPI's google/bing/youtube/… engines, a single Jules/Ollama-web-search entry — so the import step returns a usable list (source: local_catalog) instead of an error. Regression guard: tests/unit/provider-models-route.test.ts. (#5569, #5571, #5573, #5575)

  • dashboard (provider add): providers without a live key/cookie validator (e.g. LMArena (Free), PiAPI) can now be saved. The Add-connection modal treated the backend's "Provider validation not supported" response as a hard Invalid state and blocked Save entirely, leaving those providers impossible to add. The validate route now returns unsupported: true alongside the message, and the modal treats that as a non-blocking warning — the "Check" badge still shows "validation not supported" (informational), but Save persists the credential as-is. Regression guards: tests/unit/ui/add-api-key-modal-unsupported-save-5565.test.tsx (Save proceeds) and tests/unit/providers-validate-route.test.ts (wire-format). (#5565, #5567)

  • providers (codex): fix the Codex Responses WebSocket path (/v1/responses), which regressed in v3.8.40 with a client-visible Invalid JSON body and bypassed the configured proxy. (1) #5591 — PR #5237 bumped the impersonation TLS profile to chrome_149, but wreq-js@2.3.1 only supports up to chrome_147; the unknown profile produced a degenerate fingerprint and ChatGPT rejected the upstream upgrade. The Codex WS path is reverted to the proven chrome_142 (the v3.8.39 value), and the over-bumped grok-web/claude-web profiles (masked by their circuit-breaker but silently dropping TLS impersonation) are restored to chrome_146. A new regression guard asserts every configured chrome_* profile exists in the installed wreq-js typings (tests/unit/tls-profiles-valid-5591.test.mjs). (2) #5611 — the upstream wreq-js.websocket() connect ignored the Proxy Registry, so a no-direct-egress Docker container failed with a DNS error; the prepare route now resolves the Global/provider proxy and threads it through to the WS connect. Regression guard in tests/unit/responses-ws-proxy.test.mjs. (#5591, #5611)

  • providers (GLM): GLM 5.1 / 5.2 now keep the system role instead of having the system prompt folded into the first user turn. roleNormalizer.ts matched every glm* id with a blanket startsWith("glm") / startsWith("glm-") prefix, so the next-generation models — which z.ai documents as supporting the system role (GLM > 5.0) — were normalized as if they rejected it, degrading instruction-following. The matcher is now version-aware: it strips the system role only for bare glm, the 4.x family, and the 5.0 generation, and preserves it for glm-5.1/glm-5.2 (and the Fireworks glm-5p1 point alias). The ZenMux vendor-prefixed z-ai/glm-* compressed-history rule and the ERNIE rule are unchanged. Regression guards in tests/unit/role-normalizer.test.ts. (#5610)

  • Security hardening follow-ups (v3.8.15): the auth_token cookie now sets an explicit 30-day maxAge so sessions persist as intended (Seg3); the management bootstrap warns at boot when INITIAL_PASSWORD is left at the insecure CHANGEME default (Seg2); VS Code path-token endpoints (/api/v1/vscode/raw/[token]) emit a once-per-process security warning since the API key travels in the URL and can leak via logs/proxies (Seg4); the system version route resolves the real global install path via npm root -g instead of a hardcoded /app (Bug3); and auto-update mode detection segment-matches node_modules instead of substring-matching, eliminating false "global install" positives (Bug1).

  • fix(cli): rename the Node process title to omniroute so it shows correctly in ps/htop. (thanks @waguriagentic)

  • dashboard (model picker): guard against null model-alias values so opening Create Combo for a custom provider node no longer crashes. ModelSelectModal's custom-provider branch filtered modelAliases entries with a raw fullModel.startsWith(...), which threw a TypeError whenever an alias value was null/undefined (a stale/partial entry persisted to settings). The filter/map logic is extracted into a new buildNodeAliasModels helper (mirroring the sibling passthrough-alias guard, #485) that requires typeof fullModel === "string" before calling .startsWith. Regression guard: tests/unit/model-select-null-alias-guard-2247.test.ts. (thanks @wahyuzero)

  • fix(translator): strip orphaned tool results (results with no matching tool call) across request formats to avoid upstream 400s. (thanks @warelik)

  • fix(kiro): stop injecting a placeholder user turn on trailing tool-result turns so agentic loops aren't disrupted. (thanks @jetmiky)

  • fix(translator): prevent doubled tool arguments in OpenAI-to-Claude responses (duplicate finish_reason guard + string tool-input passthrough). (thanks @vishalrajv)

  • codex (agent goal streams): protect long-running agent goal streams so extended agent runs are no longer cut off prematurely. (#5772 — thanks @nguyenxvotanminh3)

  • sse (zero-width markers): strip zero-width markers from streamed responses, matching the non-streaming path so streamed output is byte-clean parity. (#5857 — thanks @DKotsyuba)

  • usage (om-usage endpoint): restore the om-usage HTTP endpoint. (#5859 — thanks @Witroch4)

  • sse (stream readiness): tune adaptive stream-readiness timeouts so slow-first-token upstreams are handled more reliably. (#5767 — thanks @nguyenxvotanminh3)

  • security (provider node URL): harden provider node URL validation. (#5760 — thanks @nguyenxvotanminh3)

  • cli (Windows doctor): correct rootDir resolution in doctor.mjs on Windows. (#5845 — thanks @arssnndr)

  • providers (Antigravity): fix a 429 hang on credit exhaustion and apply a precise reset-time model lockout instead of stalling — cleaned re-implementation of #5823. (#5846 — thanks @Chewji9875 / @diegosouzapw)

  • providers (qwen-web): unblock the validator and chat completion — the retired endpoint is replaced and the missing SPA version header is now sent. (#5855 — thanks @janeza2)

  • providers (kimi-web): migrate to the www.kimi.com Connect-RPC API after kimi.moonshot.cn was retired. (#5858 — thanks @janeza2)

  • dashboard (CSRF): unify the dashboard CSRF origin fallback so dynamic/public origins validate correctly. (#5856 — thanks @rdself)

  • db (health check interval): preserve healthCheckInterval=0 across connection create/update instead of coercing it to a default. (#5822 — thanks @atomlong)

  • sse (claude→codex streaming): stop the reasoning-summary drop and duplicated deltas on claude→codex streaming — reasoning snapshots are now synthesized in TRANSLATE mode and the sequence-number watermark is tracked per-stream (#5786). (#5832 — thanks @diegosouzapw)

  • deps (runtime): add the missing runtime dependencies @toon-format/toon and safe-regex so the published package resolves them at runtime. (#5771 — thanks @chirag127)

  • system (Windows auto-update): route in-app auto-update npm calls through the win32 shell helper so updates run correctly on Windows (#5542). (#5797 — thanks @diegosouzapw)

  • dashboard (validation badge): show a neutral badge for unsupported validation and make OAuth error messages clickable links (#5442, #5486). (#5795 — thanks @diegosouzapw)

  • providers (metadata): correct stale/broken provider metadata (#5487, #5461, #5534, #5470). (#5790 — thanks @diegosouzapw)

  • providers (local-catalog imports): import intentional local-catalog-only providers instead of surfacing a 502 (#5460, #5465). (#5787 — thanks @diegosouzapw)

  • proxyfetch (failover): skip the failover retry for non-replayable request bodies so a consumed stream isn't re-sent empty. (#5770 — thanks @Ardem2025)

  • batch (recovery): persist batch item checkpoints during recovery so an interrupted batch resumes from where it left off. (#5753 — thanks @ag-linden)

  • memory (Qdrant): enabling Qdrant now activates it as the retrieval engine (the auto default never selected it) and adds inline guidance (#5597). (#5741 — thanks @diegosouzapw)

  • chat (non-streaming aggregation): harden non-streaming SSE aggregation against malformed upstream event sequences. (#5746 — thanks @rdself)

  • sse (cooldown parsing): the anti-thundering-herd guard now tolerates numeric-epoch cooldown values. (#5747 — thanks @diegosouzapw)

  • api (body size): raise the LLM API payload limit for the responses routes so larger requests aren't rejected. (#5652 — thanks @JxnLexn)

  • providers (HuggingChat): fix HuggingChat web-session routing (#5592). (#5592 — thanks @backryun)

  • sse (heap pressure): bound the chat hot-path heap — pressure-aware admission, response cap, and clone reductions — to avoid OOM under load (#5152). (#5425 — thanks @josevictorferreira)

  • providers (M365 Copilot): validate M365 Copilot web credentials. (#5432 — thanks @skyzea1)

  • providers (chatgpt-web): restore the dot-form Pro model ids. (#5549 — thanks @Thinkscape)

  • security (error stacks): avoid rendering error stacks in responses. (#5624 — thanks @KooshaPari)

  • security (linkify): restrict linkifyText hrefs to an explicit http(s) scheme allowlist. (#948d2d7 — thanks @diegosouzapw)

  • translator (doubled tool args): prevent doubled tool-call arguments in the OpenAI→Claude translation path. (#5828 — thanks @diegosouzapw)

  • translator (orphaned tool results): strip orphaned tool-result turns across request formats so an upstream doesn't reject a tool result with no matching call. (#5805 — thanks @diegosouzapw)

  • translator (Gemini/Claude hardening): re-apply lost defensive hardening for the Gemini merge path and Claude tool defaults. (#5706 — thanks @diegosouzapw)

  • kiro (tool-result turns): stop injecting a placeholder user turn on tool-result turns, which corrupted otherwise-valid Kiro conversations. (#5807 — thanks @diegosouzapw)

  • providers (Kiro catalog): add claude-sonnet-5 to the Kiro model catalog. (#5796 — thanks @diegosouzapw)

  • oauth (connection disambiguation): disambiguate OAuth connections on username so two different identity providers no longer overwrite each other. (#5803 — thanks @diegosouzapw)

  • github (Copilot prefill): drop the trailing assistant prefill for Copilot chat, which some Copilot models rejected. (#5802 — thanks @diegosouzapw)

  • mitm (hosts cleanup): clean up privileged /etc/hosts entries on exit when possible so a crashed/interrupted run doesn't leave stale redirects behind. (#5808 — thanks @diegosouzapw)

  • dashboard (model picker): guard null modelAliases values in the model picker so a connection with no aliases no longer throws. (#5792 — thanks @diegosouzapw)

  • dashboard (error boundaries): add error boundaries for the Combos and MITM Proxy pages so a render error no longer blanks the whole dashboard. (#5788 — thanks @diegosouzapw)

  • cli (process title): rename the running process title to omniroute. (#5791 — thanks @diegosouzapw)

  • compression (context-editing telemetry): record Context Editing telemetry on the streaming path, not just the non-streaming path. (#5761 — thanks @diegosouzapw)

  • security (v3.8.15 hardening follow-ups): land the Seg2/Seg3/Seg4/Bug3 hardening follow-ups from the v3.8.15 security review. (#5512 — thanks @diegosouzapw)

📝 Maintenance
  • docs (architecture): add docs/architecture/ROUTER_BACKENDS.md — an ADR pinning down how the routing engines (ts native, bifrost, cliproxy, 9router, VibeProxy-compatible) relate to each other along two orthogonal axes (lifecycle: in-process / supervised / external vs. relay selection backend), answering the architecture questions raised in #5603 (backend interface model, why CLIProxy spawns a process, feature-flag swapping, actionable route-contract errors). The typed router-backend registry the ADR describes lands separately via #5868. (#5891)

  • tests (autoCombo): stabilize the getTaskFitnessWithSource identifies fitness_table as source for known models unit test, which flaked whenever the models.dev capabilities DB was populated in CI: the fixture model gpt-4o is a real models.dev catalog id, so the fitness resolution chain returned models_dev_tier instead of the expected static fitness_table source. The fixture now uses claude-sonnet (a shortened alias absent from the models.dev catalog, matching the sibling resolution-chain test), which deterministically falls through to the static table — the exact source and score assertions are preserved (0.95 = FITNESS_TABLE.coding["claude-sonnet"]). (#5890) — thanks @KooshaPari

  • oauth (dead-code removal): delete the superseded legacy OAuth service-class hierarchy under src/lib/oauth/services/. The live OAuth flow runs through src/lib/oauth/providers.ts + src/lib/oauth/providers/ (wired into the generic oauth/[provider]/[action] route); the old per-provider class *Service extends OAuthService implementations plus their barrel had zero production or test references. Removed oauth.ts (base class), openai.ts, github.ts, claude.ts, codex.ts, antigravity.ts, qwen.ts, qoder.ts, and the index.ts barrel (−1559 LOC). Kept the three still-live files that routes import directly by path: kiro.ts (Kiro import/exchange routes), cursor.ts (Cursor import route), and codexImport.ts (utility fns for the Codex bulk-import route). Proven safe by typecheck:core staying green (any live reference would fail the build) + a filesystem guard tests/unit/oauth-legacy-services-removed.test.ts pinning the removal against re-introduction. Salvage of the closed PR #5039. gaps v3.8.42 — T10 (5.7).

  • refactor (god-file decomposition): extracted pure leaf modules across db, sse, usage, api, memory, evals, models, resilience, and dashboard god-files (types/mappers/helpers/pure-transform leaves; behavior-preserving, test-guarded): db/providers, db/proxies, db/models, db/settings, usageAnalytics, migrationRunner (#5714, #5717, #5705, #5709, #5722, #5721); sse openai-to-gemini / cursor-protobuf / rate-limit-headers / reasoning-tag (#5824, #5794, #5736, #5734); usage families / callLogs / usageHistory / providerLimits (#5782, #5725, #5728, #5730); api provider-models discovery / unified-catalog (#5758, #5699); memory retrieval scoring (#5733); evals golden-set suites (#5740); modelsDevSync transform layer (#5743); resilience settings split (#5745); dashboard sidebarVisibility split (#5683); executor shared-utility dedup + tests (#5720 — thanks @pizzav-xyz). — thanks @diegosouzapw

  • chore (Bun script runner): adopt Bun 1.3.10 as a locked, allow-listed build/dev script runner for a small set of validated TS gate/generator scripts (Node stays the published runtime): locked runtime dependency, CI script-checks + validated-scripts run under Bun, and a bun-safe pack validator. (#5615, #5617, #5612, #5643 — thanks @KooshaPari; docs #5703 — thanks @diegosouzapw)

  • docs (sync & housekeeping): i18n CHANGELOG mirror sync for the [3.8.43] section (#5789); MCP tool count synced to 95 + routing-strategy count (#5732); README faster/leaner install notes, refreshed metrics/badges, 17-strategy + Quota-Share listing, provider counts, and grammar fixes (#5713, #5738 — thanks @chirag127); security docs for banned-keyword/account-ban detection (#5756) and the full LOCAL_ONLY route set + GHSA advisory + audit path (#5748); relay backend-routing contract clarification (#5621 — thanks @KooshaPari); release-freeze scoped to /generate-release only (#5839); .editorconfig repository standards (#5879 — thanks @shiva24082). — thanks @diegosouzapw

  • test/ci (stabilization & ratchets): guard the tsx/esm→esbuild boot transform (#5773); align t3-web web-session metadata (#5835); repoint the sidebar quota-share placement scan (#5711); lightweight health probe for batch e2e (#5651 — thanks @KooshaPari); make release-green pre-flight gates visible + bounded (#5644); stabilize nightly-mutation (tap.testFiles drift guard + anti-flake eps) (#5682); close the QG v2 tail (#5681); normalize check route paths on Windows (#5613 — thanks @KooshaPari); pass sonar.projectVersion to the SonarQube scan (#5880); plus stryker tap.testFiles registration, compression-studio smoke re-anchoring, rtk_discover de-flake, and v3.8.43-cycle ratchet rebaselines (deadExports 225→227, complexity 1981→1982, cognitive-complexity 842→845, eslintWarnings 4121→4158→4199). — thanks @diegosouzapw

  • refactor (oauth): remove dead legacy OAuth service classes. (#5838 — thanks @diegosouzapw)

🙌 Contributors

Thanks to everyone whose work landed in v3.8.43:

ContributorPRs / Issues
@ag-linden#5753
@Ardem2025#5770
@arssnndr#5845
@atomlong#5822
@backryun#5592
@baslrdirect commit / report
@Chewji9875#5563, #5579, #5846
@chirag127#5738, #5771
@DKotsyuba#5857
@hartmark#5834
@ishatiwari21#5799
@janeza2#5855, #5858
@jetmikydirect commit / report
@josevictorferreira#5425
@JxnLexn#5652
@KooshaPari#5613, #5621, #5624, #5629, #5643, #5651, #5890
@KunN-21direct commit / report
@manhdzzzdirect commit / report
@nguyenxvotanminh3#5760, #5767, #5772
@noir017direct commit / report
@pizzav-xyz#5720
@rdself#5746, #5856
@shiva24082#5879
@skyzea1#5432, #5701
@Stazyu#5557
@Thinkscape#5549
@vishalrajvdirect commit / report
@voravitldirect commit / report
@waguriagenticdirect commit / report
@wahyuzerodirect commit / report
@warelikdirect commit / report
@Witroch4#5731, #5859, #5863
@diegosouzapwmaintainer — cycle reconciliation, release-close base-red fixes, god-file decomposition, compression/memory features

What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.42...v3.8.43

View originalPermalink
How v3.8.43 went

v3.8.42

Added 5
  • Add an inflation guard to the stacked compression pipeline that discards compressed output if token count is not reduced and sends the original request instead, with a pipeline-inflation-guard warning recorded in compression stats
  • Complete German, French, and Japanese compression rule packs with dedup and ultra categories for collapsing repeated context and abbreviating technical vocabulary
  • Add Chinese (zh / wenyan) input-side compression rule pack with repeated context collapsing, pleasantry dropping, and technical term abbreviation
  • Auto-detect Chinese compression language by distinguishing zh from ja through Han-without-kana analysis
  • Add Gradle and .NET CLI (dotnet) to the RTK tool-output filter catalog with recognition of build output and preservation of build status and error messages
Fixed 2
  • Fix chatgpt-web provider failures on Electron desktop app by implementing a runtime-portable SHA-3 helper that falls back to pure-JS Keccak-f[1600] when native SHA-3 is unavailable in BoringSSL
  • Fix Bytez provider key validation by correcting the registry base URL to include the full OpenAI-compatibility chat path and replacing chat probe validation with a dedicated auth-only endpoint probe
✨ New Features
  • compression (pipeline): add an honest default-on inflation guard to the stacked compression pipeline (T02 / Headroom H1). If the fully-stacked engines produce a body that did not actually shrink — its token count is >= the original — the compressed body is discarded and the verbatim original request is sent upstream instead, with a pipeline-inflation-guard warning recorded in the compression stats. This is safe by construction (the only fallback is the unmodified original, always a valid payload) and complements the existing opt-in per-step TV1 bail-out, which governs step-to-step advancement rather than the final output. New open-sse/services/compression/pipelineGuards.ts; wired at the single finalizeStackedResult choke point shared by the sync and async stacked paths. Regression guards (incl. an inflating-engine integration test) in tests/unit/compression-pipeline-inflation-guard.test.ts.

  • compression (caveman): complete the German, French, and Japanese rule packs with the dedup (repeated-context collapsing) and ultra (abbreviation / terse) categories they were missing — these three languages previously shipped only context/filler/structural, while en/es/id/pt-BR had all five. So a de/fr/ja conversation compressed at higher intensities now collapses repeated boilerplate ("wie bereits besprochen" → "Siehe oben.", "comme mentionné précédemment" → "Voir ci-dessus.", "前述のとおり" → "(上記参照)") and abbreviates dense technical vocabulary (DatenbankDB, AuthentifizierungAuth; base de donnéesBD, authentificationauth; データベースDB, アプリケーションapp). Patterns mirror the existing es pack and stay ReDoS-safe (bounded literal alternations; the CJK pack uses no \b since Japanese has no word boundaries). Regression guard: tests/unit/caveman-packs-de-fr-ja.test.ts (packs load + validate + shrink a representative sample). gaps v3.8.42 — T05/C2.

  • compression (caveman): add a Chinese (zh / wenyan 文言) input-side rule pack — the counterpart of the existing output-side terse-cjk style. New rules/zh/{dedup,filler,ultra}.json collapse repeated context ("如前所述" → "见上。"), drop pleasantries/hedging ("请帮我…/谢谢/我觉得"), strip sentence-final modal particles ("吗/呢/吧"), and abbreviate dense technical terms ("数据库"→"DB", "应用程序"→"app"). Chinese is now auto-detected: detectCompressionLanguage distinguishes zh from ja by Han-without-kana (kana is Japanese-exclusive, so a Han-heavy Japanese sentence still resolves to ja), and zh is listed in listSupportedCompressionLanguages. Patterns are ReDoS-safe (bounded literal alternations, no \b since CJK has no word boundaries). Regression guard: tests/unit/caveman-packs-zh-wenyan.test.ts (packs load + validate + shrink; zh/ja/non-CJK detection). gaps v3.8.42 — T05/C6.

  • compression (RTK): add Gradle and .NET CLI (dotnet) to the RTK tool-output filter catalog. Tool output for gradle/gradlew and dotnet build|test|restore|publish is now recognized (both by command and by output content) and compressed: Gradle daemon/welcome banners and no-op > Task … UP-TO-DATE/SKIPPED/FROM-CACHE lines are dropped while BUILD SUCCESSFUL/FAILED, "What went wrong", and stack traces are preserved; the .NET build banner, copyright, and Determining projects to restore/Restored … chatter are dropped while Build succeeded/FAILED, error CS####/warning CS####, and test summaries are preserved. New builtin filters engines/rtk/filters/{gradle,dotnet}.json (with inline tests run by the catalog gate) plus gradle/dotnet entries in the command detector. Regression guard: tests/unit/rtk-gradle-dotnet-filters.test.ts. gaps v3.8.42 — T07/R9.

🔧 Bug Fixes
  • providers (chatgpt-web): fix 502 ChatGPT sentinel failed: Digest method not supported on the Electron desktop app, which made every chatgpt-web/* request fail. The sentinel proof-of-work hashed with native createHash("sha3-512"), but Electron's Node is built against BoringSSL, which does not implement the SHA-3 family (electron/electron#30530), so the digest threw at construction — the provider was unusable on the desktop build (works under plain Node/OpenSSL). The PoW now hashes through a new runtime-portable helper (open-sse/utils/sha3-512.ts) that prefers the native digest and transparently falls back to a dependency-free pure-JS Keccak-f[1600] when native SHA-3 is absent. The fallback is validated bit-for-bit against native createHash("sha3-512") (300 random inputs) and the published FIPS-202 known-answer vectors. Regression guards in tests/unit/chatgpt-web-sha3-boringssl-5531.test.ts. (#5531)

  • providers (bytez): fix Bytez key validation ("Provider validation endpoint not supported") and the chat base URL, verified live with a real key. Bytez is OpenAI-compatible at …/models/v2/openai/v1, but the registry stored the bare …/models/v2 base, so the validation chat-probe hit …/models/v2/chat/completions404 → the misleading "endpoint not supported". Two parts: (1) the registry baseUrl now carries the full OpenAI-compat chat path (…/models/v2/openai/v1/chat/completions); (2) key validation no longer uses a chat probe — a Bytez account only serves models explicitly added to its catalog, so even valid keys 404 on any model id. A dedicated validateBytezProvider instead probes the auth-only GET …/models/v2/list/tasks endpoint (200 ⇒ valid, 401/403 ⇒ invalid), which is independent of catalog provisioning. Regression guard: tests/unit/bytez-validation-5422.test.ts. (#5422)

  • dashboard (provider add): two provider-add UX fixes. (1) #5420 — the "Import Models" button now stays hidden for tool-only providers (web search / web fetch), not just *-search ids: firecrawl and jina-reader (declared serviceKinds: ["webFetch"]) previously showed an Import button that hit the 400 "does not support models listing" route. A new capability check (providerLacksModelListing over the resolved serviceKinds) gates the section without ever hiding an LLM/media provider. (2) #5426 — Coze key validation no longer leaks the raw upstream envelope ({code,msg,logId,from}) into the UI; the Coze-shaped error becomes a friendly Coze rejected the key: <msg> (code <n>) message (scoped to provider === "coze" so no other provider is affected). Regression guards: tests/unit/model-listing-capability-5420.test.ts, tests/unit/coze-validation-error-5426.test.ts. (#5420, #5426)

  • providers (friendliai, novita): fix two provider registry endpoints that rejected valid keys (verified live with real keys). FriendliAI pointed at …/dedicated/v1/chat/completions, which 403 Forbiddens a serverless flp_* token — switched to …/serverless/v1/chat/completions (+ a serverless modelsUrl). Novita pointed at the legacy …/v3/… base with a typo'd model id ai-ai/llama-3.1-8b-instruct (both 404) — switched to the OpenAI-compatible …/openai/v1/… base + the valid meta-llama/llama-3.1-8b-instruct id. Regression guard: tests/unit/provider-endpoints-friendliai-novita.test.ts. (#5430, #5455)

  • providers (muse-spark): align the Muse Spark Web (Meta AI) cookie copy with the live cookie name. The default session cookie migrated from the retired abra_sess to ecto_1_sess (META_AI_DEFAULT_COOKIE), but the provider form hint and one 401 auth-failure message still told users to paste abra_sess — a cookie that no longer exists. Both strings now name ecto_1_sess. Regression guard: tests/unit/muse-spark-cookie-copy-5449.test.ts. (#5449)

  • dashboard (provider add): fix three rough edges in the Add-API-Key / model-import flow reported across the provider-catalog audit. (1) The Validation Model and Account ID form fields shipped untranslated i18n stub copy ("Validation Model Id Label", "Account Id Placeholder", …) that surfaced verbatim in the modal — replaced with real labels/placeholders/hints in en.json. (2) Model import silently fell back to the cached/local catalog: the route already returned a warning ("API unavailable — using local catalog"), but useModelImportHandlers only read models/error and dropped it, so the user got local models with no indication — the warning is now surfaced as an import log line (new pure helper extractImportWarning). (3) The required connection-name field defaulted to "", which let browser autofill inject garbage (e.g. wiw) — it now defaults to "main". Regression guard: tests/unit/provider-add-ux-i18n-import-warning.test.ts. (#5421, #5428, #5429, #5431, #5435)

  • services (installer): fix spawn EINVAL when installing an embedded service (9Router / CLIProxy) on Windows + Node.js 24+. Node 24 stopped letting child_process.execFile() run .cmd batch files without a shell (nodejs/node#52554), and npm on Windows is npm.cmd, so runNpm() threw EINVAL the moment a user clicked Install. runNpm now enables shell on win32 only. To keep Hard Rule #13 intact under a shell — where the shell, not execFile, parses argv — the install --prefix (a DATA_DIR path that can legitimately contain spaces, e.g. C:\Users\John Doe\.omniroute\…) is now passed via the npm_config_prefix environment variable instead of an argv path, and the user-supplied install version is constrained to a dist-tag/semver shape (SERVICE_VERSION_PATTERN) at the route boundary so it can never carry shell metacharacters. With the prefix in the environment and the version validated, every remaining argv entry is a static flag. Regression guards: tests/unit/services/installers/runNpm-shell-5379.test.ts (+ existing ninerouter.test.ts aligned to npm's npm_config_prefix env). (#5379)

  • cli (serve): restore dist/tls-options.mjs to the npm tarball — the opt-in native HTTPS/TLS sidecar (#5361) was copied into the staged dist/ by the build but then pruned by the prepublish allowlist step, so omniroute serve crashed on the published 3.8.41 with ERR_MODULE_NOT_FOUND (dist/server-ws.mjs imports ./tls-options.mjs). Added tls-options.mjs to APP_STAGING_ALLOWED_EXACT_PATHS (survives the prune) and dist/tls-options.mjs to PACK_ARTIFACT_REQUIRED_PATHS (the check:pack-artifact gate now fails loudly if it ever vanishes again — same guard pattern as webdav-handler.mjs). Regression guards in tests/unit/pack-artifact-policy.test.ts. (#5452 — thanks @KooshaPari for the parallel fix #5494)

  • dashboard: fix the Add Provider / onboarding wizard button silently doing nothing. The /dashboard/providers/new route was a redirect stub (it bounced straight back to /dashboard/providers), so every "Add Provider" button and dashboard widget link opened nothing, and the fully-built ProviderOnboardingWizard component stayed orphaned (never rendered by any route). The route now renders the wizard directly; auth is enforced centrally by the (dashboard) layout, same as the sibling provider routes. Regression guard in tests/unit/onboarding-wizard-route-5427.test.ts. (#5427)

  • db (import): fix EBUSY: resource busy or locked when importing a database on Windows. The import route deleted the live storage.sqlite + WAL/-shm/-journal sidecars with a plain fs.unlinkSync immediately after resetDbInstance(), but Windows releases the SQLite file handle asynchronously after close() (mmap / antivirus), so the unlink raced and threw EBUSY. The route now deletes via unlinkFileWithRetry (EBUSY/EPERM backoff) — the same helper the restore path already uses. Regression guard in tests/unit/db-import-ebusy-5406.test.ts. (#5406, consolidated under #5161)

  • build: keep ioredis out of the client/CLI bundle — a dast-smoke regression revealed the module was being pulled into browser/Electron client-side chunks; adding it to the SPAWN_CAPABLE_PREFIXES leaf excludes it from client bundles while keeping it available on the server path. (#5546)

  • providers (mimocode): route per-account traffic through SOCKS5 proxy dispatchers — each mimocode account's requests are now dispatched via its configured SOCKS5 proxy rather than the default direct connection. (#5521 — thanks @pizzav-xyz)

  • providers: persist the Configured provider filter selection across page reloads — the filter was resetting to "All" on every navigation. (#5510 — thanks @KooshaPari)

  • providers (chatgpt-web): support GPT-5.5 Pro model handoff — adds the model mapping and handoff routing needed for the GPT-5.5 Pro tier. (#5536 — thanks @Thinkscape)

  • dashboard: keep onboarding schemas browser-safe — the schema module imported a server-side db reference that crashed the browser bundle; it is now imported only on the server path. (#5525 — thanks @KooshaPari)

  • routing (bifrost): add auto-fallback cooldown for bifrost targets — prevents rapid re-selection of a failing bifrost backend within the cooldown window, complementing the existing circuit-breaker mechanism. (#5519 — thanks @KooshaPari)

  • providers (opencode-plugin): bump the opencode plugin to v0.2.0 and wire auto-publish on release so the plugin package tracks OmniRoute releases automatically. (#5363 — thanks @herjarsa)

  • rate-limit: normalize queue refresh settings — aligns the queue-refresh interval configuration across rate-limit strategies so stale queues are released on a consistent schedule. (#5499 — thanks @KooshaPari)

  • fallback: normalize provider error-rule header extraction — ensures fallback retry decisions correctly read all response headers regardless of casing, fixing cases where a provider's Retry-After or custom error header was silently dropped. (#5473 — thanks @KooshaPari)

  • routing: gate Claude adaptive-thinking defaults behind the feature flag — prevents the thinking budget from being injected into requests for models that do not support the extended-thinking parameter, avoiding upstream 400 errors on non-thinking Claude variants. (#5480 — thanks @KooshaPari)

  • ci: fix post-merge CI regressions introduced by the dead-code sweep — restores test imports and type references broken when the ratchet landed before downstream consumers were updated. (#5467 — thanks @KooshaPari)

  • sse: treat terminal stream cancels as complete — an aborted SSE stream was being left in a partial state, causing downstream consumers to wait indefinitely for a final event that would never arrive. (#5491 — thanks @JxnLexn)

  • api: fix framing of non-streaming JSON responses — stream: false chat-completions responses were returned without correct content-length framing, causing some clients to misparse the response body. (#5416 — thanks @rdself)

  • dashboard (tests): protect dynamic dashboard endpoint tests with CSRF validation — the test suite was exercising dashboard API routes without CSRF tokens, masking a coverage gap for those endpoints. (#5405 — thanks @rdself)

  • providers: remove the dead Phind provider (service shut down) and deduplicate the HuggingChat catalog listing that had accumulated a stale duplicate entry. (#5530 — thanks @backryun)

  • providers (longcat): correct the LongCat free tier — LongCat-2.0 is now GA; the one-time 10M-token promo (KYC required) is correctly reflected in the catalog, replacing the stale legacy beta entry. (#5508 — thanks @backryun)

📝 Maintenance
  • dashboard (refactor): consolidate the duplicate caveman on/off toggle from the compression settings tab onto the single-source panel (T11), eliminating the stale off-sync copy. (#5524)

  • tests: add quota guard for Claude-Code identity version lockstep (Phase 2) — asserts that the Claude-Code version reported in quota accounting stays in sync with the deployed version, preventing silent drift. (#5514)

  • docs: add relay backend strategy guide documenting supported relay backend types, selection criteria, and configuration patterns. (#5547)

  • docs: clarify bifrost relay backend environment variables — documents which env vars control bifrost's relay backend selection and failover behavior. (#5520 — thanks @KooshaPari)

  • tests: add relay routing fallback header behavior tests — regression guard asserting that fallback-triggered relay requests carry the correct forwarded headers through the routing layer. (#5526 — thanks @KooshaPari)

  • ci: add npm fetch-retry configuration and codify the release-freeze protocol (Hard Rule #21) — reduces transient npm registry fetch failures in CI and establishes the documented procedure for freezing releases. (#5506)

  • deps: bump 11 production dependencies to their latest compatible versions. (#5414)

  • deps: bump Electron from 42.4.1 to 42.5.1 in /electron. (#5413)

  • deps: bump the development dependency group with 9 updates. (#5415)

  • maintenance (dead-code): repo-wide sweep of unused exported symbols, types, and schemas — removes 35 no-longer-referenced exports across cloud-agent, a2a, SSE, memory, quota, skills, gamification, codex, qdrant, playground, provider catalog, and combo modules, reducing the exported API surface and eliminating stale misleading types. (#5372, #5373, #5374, #5375, #5376, #5377, #5378, #5380, #5381, #5382, #5383, #5384, #5385, #5386, #5387, #5388, #5389, #5390, #5391, #5392, #5393, #5395, #5396, #5397, #5398, #5399, #5400, #5401, #5402, #5403, #5404, #5463, #5464, #5466, #5468 — thanks @JxnLexn)

  • maintenance (DRY): DRY consolidation of shared helpers — extracts 17 duplicated utilities into single shared modules: vscode metadata helpers, proxy route handlers, auth zip extractors, combo-builder model options, vscode tokenized-request helpers, quota strategy ranking helpers, recharts donut card, provider-specific validation, batch response formatter, Redis runtime helpers, version-manager request parsing, media-generation route helpers, service install helpers, settings transform schemas, relay stream finalizer, machine-id fallback, and node SQLite adapter. (#5471, #5472, #5475, #5477, #5479, #5482, #5484, #5485, #5488, #5490, #5492, #5493, #5495, #5496, #5497, #5498, #5500 — thanks @JxnLexn)


What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.41...v3.8.42

View originalPermalink
How v3.8.42 went

v3.8.41

Added 4
  • Add selectable relay backend (TS / Bifrost / auto) via OMNIROUTE_RELAY_BACKEND / RELAY_ROUTING_BACKEND environment variable, with auto mode selecting Bifrost when BIFROST_BASE_URL is set and falling back to TypeScript if the sidecar is unreachable
  • Add X-Routing-Backend and X-Routing-Fallback response headers to indicate which relay backend handled the request
  • Add Proxy Pool dropdown to OpenCode Free per-account proxy modal, allowing selection of pre-saved proxies instead of manual Host/Port/credentials entry
  • Add Saved / Custom toggle for proxy configuration in OpenCode, with server-side resolution of proxies from the registry
Fixed 12
  • Fix Claude translator to synthesize a minimal user turn when an OpenAI request carries only system/developer messages, preventing 400 errors from the Messages API
  • Remove retired Google AI Studio model IDs and align the Gemini catalog to live GenAI API, removing gemini-1.5-pro, gemini-1.5-flash, gemini-2.0-flash, gemini-2.0-flash-lite, and adding gemini-3.1-flash-lite and gemini-embedding-001/gemini-embedding-2
  • Add modelDeprecation forwards for renamed Gemini model IDs to redirect legacy IDs to GA models instead of returning 404
  • Fix embedded-services dashboard by lazily initializing service supervisors from /api/services/[name]/logs to prevent 404 errors for cliproxy and 9router logs before bootstrap
  • Fix dashboard lifecycle buttons to send JSON and properly handle empty install bodies defaulting to version latest
  • Fix dashboard lifecycle and log-stream failures to surface as actionable UI errors instead of silently showing no logs
[3.8.41] — 2026-06-29
✨ New Features
  • feat(relay): selectable relay backend (TS / Bifrost / auto) — the OpenAI-compatible relay endpoint can now route its hot path through a native Bifrost sidecar without clients changing URLs. OMNIROUTE_RELAY_BACKEND / RELAY_ROUTING_BACKEND = ts | bifrost | auto: defaults to the existing TypeScript relay; auto selects Bifrost when BIFROST_BASE_URL is set (and BIFROST_ENABLED0) and falls back to TS automatically if the sidecar is unreachable; bifrost keeps strict failure behavior. Auth, per-IP/token rate limits, prompt-injection checks, and model allowlists still run in the Next relay route before dispatch (control plane stays in the app); responses carry X-Routing-Backend / X-Routing-Fallback. Regression guards: tests/unit/api/v1/relay-routing-backend.test.ts, tests/unit/api/v1/bifrost-sidecar.test.ts. (#5315, #5316 — thanks @KooshaPari)
🔧 Bug Fixes
  • translator (claude): synthesize a minimal user turn when an OpenAI→Claude request carries only system/developer messages, so the request stops failing with [400]: messages: at least one message is required. openaiToClaudeRequest hoists every system/developer turn into Claude's top-level system field and filters them out of messages; an all-system input (OpenCode compaction / title-generation requests) left messages: [], which the Messages API rejects — surfacing in OpenCode as a mid-task stream error that drops the conversation. The guard fires only when messages would otherwise be empty (system instructions still drive the response), so non-empty requests are unaffected. (#5342 — thanks @wild-feather)
  • providers (gemini): drop retired Google AI Studio model ids and align the catalog to what the live GenAI API actually serves (verified 2026-06-29 against the official deprecations page). Removes long-retired gemini-1.5-pro/gemini-1.5-flash, the shut-down gemini-2.0-flash/gemini-2.0-flash-lite, and dead experimentals; renames gemini-3.1-flash-lite-preview → the GA gemini-3.1-flash-lite; swaps the retired text-embedding-004 for the live gemini-embedding-001/gemini-embedding-2; and adds graceful modelDeprecation forwards so legacy/renamed ids redirect to the GA model instead of 404ing. Native AI-Studio-direct image/video/music registration is intentionally out of scope (needs real executor work; those models stay reachable via Antigravity/Vertex/aggregators). (#5337 — thanks @backryun)
  • services (dashboard): fix the embedded-services dashboard failures (#5298) — service supervisors are now lazily initialized from /api/services/[name]/logs so cliproxy/9router logs no longer 404 before bootstrap registers a supervisor; lifecycle buttons send JSON (empty install bodies default to version: "latest", malformed JSON still returns 400 Invalid JSON body); lifecycle and log-stream failures surface as actionable UI errors instead of silently showing no logs; Tailscale CGNAT 100.64.0.0/10 peers count as private-LAN local for local-only service access; a parent /dashboard/context/dashboard/context/settings redirect stops RSC prefetch 404s; and /api/v1/providers/{cliproxyapi,9router}/models return synced embedded-service models instead of invalid_provider. (#5299, #5298 — thanks @KooshaPari)
  • thinking (claude): fix three independent defects in Claude adaptive-thinking on the OpenAI-compatible path (Cursor → Claude OAuth). (A) the dashboard Thinking-Budget setting was dropped on every restart — setThinkingBudgetConfig was never called at boot, so a saved {mode:"adaptive"…} silently reverted to passthrough; it's now hydrated from settings in server-init. (B) the Claude executor force-injected adaptive thinking after translation, ignoring the operator's budget — it now honors mode:"auto" (strip) while keeping the default (passthrough) behavior byte-identical so native Claude Code is unaffected, and remaps an operator thinking.type:"enabled" to the adaptive shape Opus 4.7/4.8 require (enabled → 400). (D) on replay, signature-less reasoning_content was reconstructed as a thinking block carrying a fabricated signature → Anthropic 400 "Invalid signature in thinking block"; it now emits a signature-less redacted_thinking block (real signatures are still preserved verbatim). Regression guards: tests/unit/thinking-budget-hydration-5312.test.ts, base-thinking-budget-config-5312.test.ts, openai-to-claude-redacted-replay-5312.test.ts (existing #5123/#4479/#2454 suites stay green). The </think> content-marker channel mismatch (RC-C, shared with #5245) is tracked as a follow-up pending a live Anthropic validation. (#5312 — thanks @vitalNohj)
  • opencode (proxy pool): the OpenCode Free per-account proxy modal now offers the global Proxy Pool dropdown (by-id reference) instead of forcing manual Host/Port/credentials on every account — Gap 1 of #5217. A Saved / Custom toggle: "Saved" picks a pre-saved proxy from GET /api/settings/proxies and stores {fingerprint, proxyId}, so updating that pool proxy applies to every account using it; "Custom" keeps the manual inputs (stored inline) as an escape hatch. Resolution happens server-side (resolveAccountProxiesFromRegistry) so the executor still receives a resolved proxy unchanged; existing inline entries keep working and an unknown/deleted proxyId degrades safely to direct. Regression guards: tests/unit/noauth-proxy-resolution.test.ts, tests/unit/ui/noauth-account-card.test.tsx. (#5217 Gap 1 — thanks @daniij)
  • thinking (claude): let reasoning_content-native clients (e.g. Cursor) opt out of the </think> close-marker so it no longer leaks an orphan </think> into visible content (RC-C of #5312, shared with #5245). The marker-suppression machinery already existed (UA allowlist, #5348) but Cursor's UA was deliberately excluded; this adds an explicit request header x-omniroute-thinking-marker: off (also on/keep to force-keep) that overrides the UA policy. With the header absent the behavior is byte-identical — Claude Code/Cursor-composer clients that scan content for the marker (#4633) still receive it. Regression guard: tests/unit/think-close-marker-suppress-5245.test.ts (#5123 case-b + #4479 stay green). (#5312, #5245 — thanks @vitalNohj, @wild-feather)
  • cors: browser/Electron clients (e.g. Wayland AI) can now use OmniRoute as an OpenAI-compatible provider out-of-the-box. The token-authenticated API surface (/v1/*, /v1beta/*) now returns a permissive Access-Control-Allow-Origin (echoes the request Origin, * when absent) by default — matching 9router and the OpenAI-compatible ecosystem — so a renderer fetch can read the response instead of failing CORS-blocked as "site not found" / empty catalog (while curl, which sends no preflight, worked). This is safe: those routes auth via Authorization/x-api-key headers browsers never auto-attach (no credentialed-session/CSRF exposure), and Access-Control-Allow-Credentials is never paired with the echo/wildcard. Cookie-authed MANAGEMENT/dashboard routes stay exactly fail-closed; CORS_ALLOW_ALL/CORS_ALLOWED_ORIGINS still take precedence. Regression guards: tests/unit/cors/origins.test.ts, tests/unit/authz/pipeline.test.ts. (Bug 2 of #5242 — thanks @jonlwheat2-gif)
  • grok-web: forward the Cloudflare clearance cookies and stop mislabeling IP-reputation blocks as a bad cookie. "Check cookie" returned Invalid SSO cookie even with a valid, complete browser session — but the cookie parser was never the problem (it robustly extracts sso/sso-rw from a full DevTools header). Two real gaps fixed: (1) buildGrokCookieHeader now forwards cf_clearance and __cf_bm when pasted (it dropped them before; AIClient2API forwards them too) — strictly additive, a bare sso blob still yields exactly sso=…; (2) when the user supplied a cf_clearance, a 401 / invalid-credentials-403 from grok.com is now surfaced as an IP-reputation/anti-bot block (cf_clearance is IP+TLS+UA-pinned and can't be replayed from a different machine) instead of the misleading "Invalid SSO cookie — re-paste". A bare cookie with no clearance still gets the re-paste hint. Regression guards in web-cookie-auth.test.ts + provider-validation-specialty.test.ts. (#5350 — thanks @SeaXen)
  • cli (serve): opt-in native HTTPS/TLS for omniroute serve — so strict-CSP Electron apps and browsers can reach OmniRoute over https:// instead of plain http://localhost. Provide --tls-cert <path> --tls-key <path> (or OMNIROUTE_TLS_CERT/OMNIROUTE_TLS_KEY) and the standalone server terminates TLS on the same listener (no extra port/proxy); WebSocket upgrade (live dashboard + /v1 streaming) works over wss:// unchanged since https.Server extends http.Server. With no TLS flags the HTTP path is byte-identical to before; only one of cert/key, or an unreadable path, logs a warning and stays HTTP (never half-enables, never crashes). Auto-generated self-signed certs for localhost are a follow-up; for now provide an explicit cert/key (or front OmniRoute with a TLS terminator). Regression guard: tests/unit/tls-options.test.ts. (Bug 1C of #5242 — thanks @jonlwheat2-gif)
  • opencode/observability: make OpenCode Free account/proxy rotation visible and fix two real defects surfaced alongside it. (1) the per-request rotation selection log (dispatch via account … through proxy …) was debug (hidden at default APP_LOG_LEVEL=info) — promoted to info so the shuffle/cooldown lifecycle is auditable (token stays masked). (2) [ProxyEgress] reported proxy=direct even when an account proxy was applied, because the egress logger ran outside the executor's nested proxy context — the effective applied proxy is now captured (via an applied-proxy sink threaded through the proxy AsyncLocalStorage) and reflected in the egress log. (3) [callLogs] too many SQL variablesdeleteCallLogRowsByIds deleted up to 5000 ids in one IN (…), exceeding SQLite's ~999 bound-param cap and aborting log trimming/retention; ids are now chunked (≤500 per statement). Regression guards: tests/unit/call-log-trim-sql-vars-5217.test.ts, apply-executor-proxy-info-5217.test.ts, extended opencode-proxy-rotation-4954.test.ts. The Proxy Pool dropdown (by-id) UI (Gap 1) is a follow-up requiring browser validation. (#5217 — thanks @daniij)
  • chatgpt-web: wire tool/function calling into the chatgpt-web provider. It was the only web-session executor that never read body.tools — both response builders hardcoded finish_reason:"stop" and emitted only content, so tool calls were silently dropped (the model answered in prose). It now uses the shared webTools prompt-emulation shim (a <tool>-contract system message + <tool>{…}</tool> response parsing) exactly like its 9 sibling executors (qwen-web, perplexity-web, …) — it was simply omitted from the #3259 rollout. Tool mode buffers and emits tool_calls + finish_reason:"tool_calls" (gated off the image-gen path); plain chat is unchanged. Regression guard: tests/unit/chatgpt-web-tools-5240.test.ts. (#5240 — thanks @Rougler)
  • oauth/dashboard: fix the persistent/false Antigravity "Token Expired" badge (continuation of #3679/#3850). Two causes: (1) new OAuth connections never set tokenExpiresAt (only expiresAt), so the dashboard badge — which prefers tokenExpiresAt || expiresAt — fell back to the original grant clock and could flash a false "Token Expired" until the first background refresh. Creation now mirrors expiresAt into tokenExpiresAt across all 5 OAuth create paths (a shared buildOAuthConnectionCreatePayload), consistent with every refresh path which already writes both. (2) when a refresh-capable connection has no usable refresh token, the health-check sweep silently skipped it, leaving testStatus="active" forever while the cosmetic badge showed expired; it now surfaces a terminal testStatus="expired" ("needs re-auth"), tightly gated so it never clobbers non-refresh providers or already-terminal/cooldown states. Regression guards: tests/unit/oauth-connection-tokenexpiresat-5326.test.ts, tests/unit/token-health-no-refresh-token-expired-5326.test.ts. (#5326)
  • routing: auto-disable a depleted API key on upstream 402 "Insufficient account balance" for API Key Round-Robin connections (multiple keys in one connection's extraApiKeys). The per-connection path already terminalized 402 (→ credits_exhausted), but the per-KEY health tracker (recordKeyHealthStatus) only recorded failures for 401, so a 402-depleted key stayed in rotation and kept getting retried. Now a 402 marks the current key invalid immediately (terminal — balance won't recover mid-session) via a new recordKeyTerminal, so the rotator skips it and falls over to the next healthy key; the state persists across restarts. Also added insufficient balance/insufficient_balance/insufficient account balance to the credits-exhausted body signals so non-402 out-of-credit responses terminalize too. Regression guard: tests/unit/key-health-402-disable-5239.test.ts. (#5239 — thanks @muflifadla38)
  • cli: omniroute serve no longer discards a user-set NODE_OPTIONS=--max-old-space-size=…. It used to unconditionally overwrite NODE_OPTIONS (and pass an explicit --max-old-space-size CLI arg) with the calibrated default, so a user who exported --max-old-space-size=8192 still ran at the old cap and OOM'd (#5238 reporter set 8192, crashed at ~505MB). Now it mirrors the Electron and standalone launchers: if NODE_OPTIONS already pins the heap, that value wins (and the duplicate CLI arg is suppressed); otherwise the calibrated --max-old-space-size is appended, preserving unrelated flags. Regression guard: tests/unit/serve-node-options-preserve-5238.test.ts. (Defect C of #5238; the b.mask/OOM-root parts are tracked separately.)
  • dashboard: restore the {active}/{total} active model-count badge in a provider's Available Models toolbar (provider detail page). It was dropped during the v3.8.13 god-file decomposition (#3327) — the ModelVisibilityToolbar still received activeCount/totalCount but they were orphaned as unused _-prefixed params and the rendering <span> was never carried over (the modelsActiveCount i18n key stayed). Re-wired the existing props to the existing key; zero data-layer or i18n change. Regression guard: modelVisibilityToolbarActiveCount.test.tsx. (#5264)
  • rerank: /v1/rerank no longer rejects SiliconFlow and DeepInfra Qwen3-Reranker models with 400 "Invalid rerank model" even though /v1/models lists them. The model-ID parser was never the problem (it already splits on the first slash, so siliconflow/Qwen/Qwen3-Reranker-8B parses correctly) — siliconflow and deepinfra were just missing from the rerank provider registry. Added both: SiliconFlow as Cohere-compatible, DeepInfra via a new deepinfra adapter (model in the URL path POST /v1/inference/<model>, {queries,documents} request, positional {scores} response mapped to Cohere results[]). Regression guard: tests/unit/rerank-providers-5332.test.ts. (#5332 — thanks @maikokan)
  • authz/dashboard: stop rejecting every dashboard mutation with 403 INVALID_ORIGIN when the dashboard is reached over a LAN IP / non-localhost host. The origin-pinning check (#5278) only accepted the configured *_PUBLIC_BASE_URL (typically http://localhost:20128) plus the internal request.url origin — which Next.js standalone reports as the bind host, not the real Host. So opening the dashboard at e.g. http://192.168.0.15:20128 made the browser's same-origin Origin match no candidate, and every POST/PUT/DELETE (save API key, save provider, test connection) failed while GETs still worked. Two fixes: (a) the request Host (or a trusted X-Forwarded-Host) is now accepted as a valid mutation origin, gated by two independent checks — the token-stamped socket peer must be loopback/private-LAN and the Host itself must be a loopback/private-LAN IP literal, so a DNS-rebinding domain (which classifies as remote) can never become a trusted origin and the protocol is pinned to the actual connection; (b) the INVALID_ORIGIN response now carries an actionable message (set OMNIROUTE_PUBLIC_BASE_URL) and the dashboard surfaces API error .message via a shared extractApiErrorMessage helper instead of rendering the raw error object. Regression guards: tests/unit/authz/public-origin.test.ts (direct LAN/loopback + DNS-rebinding defense), tests/unit/api-error-message-5340.test.ts. (#5340)
📝 Maintenance
  • chore(dead-code): repo-wide sweep of unused exported symbols and a matching dead-code baseline ratchet — trimmed unused exported helpers, validation/settings/encryption-config schemas, utility/domain/static-constant/formatting helpers, runtime test helpers, the request-timeout fetch wrapper, event-bus, semantic-cache (maintenance + expiry), correlation-middleware, MCP-scope, service-registry, build-profile, api-key-format, authz-class, models.dev-context, embedding-cache, provider-limits-scheduler, search-validator, webhook-example, agent-skills-repo-URL and command-code-auth-cleanup exports. Pure dead-code removal validated by typecheck:core (no remaining referencing site) — no behavior change. (#5321, #5322, #5324, #5325, #5328, #5329, #5330, #5331, #5333, #5334, #5335, #5336, #5338, #5339, #5353, #5354, #5355, #5356, #5357, #5359, #5362, #5364, #5365, #5366, #5368, #5369, #5371 — thanks @JxnLexn)

What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.40...v3.8.41

View originalPermalink
How v3.8.41 went

v3.8.40

Added 8
  • Add relevance extractive compression engine that scores sentences by term-overlap with user queries and greedily keeps the most relevant within a budget
  • Add hard-budget compression mode that trims bodies to a token budget by ranking and dropping lowest-saliency sentences while preserving original order
  • Add opt-in result memoization for deterministic compression engines to skip recompute on the hot path
  • Extend X-OmniRoute-Compression response header to surface compression transparency annotation with token counts and rules applied
  • Add saliency heatmap visualization in the compression studio to color tokens by saliency score
  • Add composite-command splitter for RTK detection to recognize individual commands in chains like cd /x && git status
  • Add omniroute_tool_search MCP tool for lexical keyword search over MCP tool names and descriptions
  • Add opt-in RTK semantic command-output renderers that rewrite structured command output into compact forms
Changed 1
  • Extend X-OmniRoute-Compression header to append compression statistics while maintaining backward compatibility with existing parsers
[3.8.40] — TBD

In development — bullets added per PR; finalized at release.

✨ New Features
  • feat(compression): relevance extractive engine — a new opt-in compression engine that scores each sentence by term-overlap (Jaccard) with the user's last query minus a length/boilerplate penalty, greedily keeps the most relevant within a budget, and reconstructs the original order. Pure-string, deterministic, ReDoS-safe (char-code tokenization, no RegExp over user input), fail-open, default off. Ideal for trimming long pasted RAG context / tool output to what's relevant. Sentences carrying real signal (digits/URLs/errors/code/paths) are never dropped; overlapThreshold/budgetPercent/boilerplateWeight are configurable. Tier-2 item of the compression feature-extraction roadmap (#7). (#5289)
  • feat(compression): hard-budget mode — compress to ≤ N tokens — a deterministic post-pass (targetTokens / targetRatio, default unset → no-op) that trims a body to a token budget. It ranks sentences/lines by average scoreToken ascending and drops the lowest-saliency ones until the body fits (measured by the exact cl100k countTextTokens), preserving original order. Lines carrying real signal (digits, URLs, Error:-family, code fences, stack at-frames, multi-segment paths, key=value) are never dropped; the budget is distributed proportionally across messages so the total stays ≤ target; an unreachable target (all-preserved) surfaces a validationWarnings note instead of failing silently. Does NOT touch the estimateCompressionTokens budget-gate estimator. Tier-3 item of the compression feature-extraction roadmap (#17). (#5288, follow-up #5291)
  • feat(compression): result memoization for deterministic engines (opt-in) — caches (input, config) → result for provably pure, stateless modes (lite/standard/rtk and stacked pipelines of {lite,caveman,rtk}) to skip recompute on the hot path. Opt-in via memoizeCompressionResults (default off → zero behavior change). Conservative opt-in whitelist (stateful ccr/session-dedup — which write the cross-request CCR store — and model-backed ultra/aggressive/llmlingua are never cached), principal-scoped (skipped without a principal, so no cross-principal body leak), and clone-on-store + clone-on-read. Tier-3 item of the compression feature-extraction roadmap (#21). (#5286)
  • feat(compression): inline transparency annotation — surfaces tokens=847→312; rules: filler×8, dedup×2 derived from existing compression stats. The X-OmniRoute-Compression response header is extended append-only (the mode; source=X prefix stays byte-identical, so existing header parsers don't break) and the compression studio cockpit shows a matching badge. Zero new computation — it aggregates the rulesApplied/techniquesUsed already on the stats. Tier-3 item of the compression feature-extraction roadmap (#18). (#5284)
  • feat(compression): saliency heatmap in the compression studio — the preview studio can now color each token by saliency: ultra per-token scoreToken (0–1, green→red gradient) or universal kept/removed from the existing diff. A dry-run visualization behind a toggle (no cost on a normal preview; backward-compatible when off). Completes the visualization half of roadmap item #13 (the A/B comparison shipped in #5080). (#5285)
  • feat(compression): composite-command splitter for RTK detectioncd /x && git status now detects as git-status (previously the whole string was treated as one command and matched no filter). A quote-aware top-level tokenizer splits on &&/||/; (never inside quotes or $(…)/backtick subshells) and feeds the last segment to RTK command detection, so every RTK filter/renderer fires on commands wrapped in cd … &&/||/; chains. O(n), no RegExp over the command (ReDoS-safe). Tier-3 item of the compression feature-extraction roadmap (#16). (#5283)
  • feat(mcp): omniroute_tool_search tool + one-line TS signatures — new MCP tool that does lexical keyword search over every MCP tool's name/description and returns the top matches as compact one-line TypeScript signatures (~half the JSON-schema token cost), so agents discover tools on demand instead of carrying all ~88 schemas every turn. Search is ReDoS-safe (substring scoring, never new RegExp on the query) and deterministic; tools/list stays complete (no hidden tools). Adds the read:tools scope. Tier-1 item of the compression feature-extraction roadmap. (#5269)
  • feat(compression): RTK semantic command-output renderers (opt-in) — adds a second, opt-in compaction layer to the RTK engine that rewrites structured command output into a far more compact semantic form: git diff → file headers + @@ hunks + changed lines only; an all-green pytest/jest/vitest/eslint run → its one-line summary; terraform/tofu planPlan: +N ~M -K plus the resource list; kubectl/aws JSON arrays → a minimal table. Each renderer is conservative (no-op when the shape doesn't match) and the integration is fail-open; the test-green renderer never collapses output that carries any failure signal. Gated by RtkConfig.enableRenderers (default off → zero behavioral change). Eighth item of the compression feature-extraction roadmap. (#5268)
  • feat(compression): QuantumLock cache-prefix stabilization (opt-in, default off) — recovers upstream prompt-cache hits that a volatile fragment in the system prompt would otherwise bust. When a caller injects a session UUID, unix timestamp, request-id, JWT, API-key shape, or long hex digest into the role:system message every turn, the longest common prefix across turns ends at that changing byte → the whole system prompt after it is re-billed and re-processed each turn. QuantumLock replaces each non-semantic volatile fragment with a positional, value-independent placeholder ⟦Q{i}⟧ and appends the real values in a delimited ⟦QUANTUMLOCK⟧ tail. The rewrite is sent to the model (lossless — not restored), so the system-prompt body becomes byte-identical across turns and the provider caches the long stable prefix while only the small tail differs. Opt-in, default off, applied only for caching providers (isCachingProvider && config.quantumLock.enabled); bounded ReDoS-safe patterns; idempotent; no date/time patterns (semantically meaningful — explicit non-goal). Studio gets a toggle + a "🔒 N volatile fragment(s) stabilized" dry-run badge. Seventh item of the compression feature-extraction roadmap (bench: #5080, gate: #5127, fuzzy: #5143, ionizer: #5148, TOON: #5163, CCR ranged: #5187, risk-gate: #5243). (#5260)
  • kilocode: anonymous (no-auth) access to Kilo Code's free models, mirroring the opencode/mimocode pattern. With no Kilo account connected, requests now fall back to the gateway's anonymous tier (Authorization: Bearer anonymous on api.kilo.ai/api/openrouter) so the free models work without signup; a connected OAuth account is still used unchanged for the paid tier (#5259, #4019 — thanks @Theadd for the reference implementation)
  • feat(logging): call-log correlation ID (end-to-end) — every request now gets a unique correlation id, returned in the X-Correlation-Id response header, persisted in call_logs (migration 109), filterable via /api/usage/call-logs, and surfaced in the dashboard request logger (per-chunk stream timestamps + active-requests-first sort). This is the safe, cohesive core subset of the larger #5275 — landed on its own so the low-risk value isn't blocked by the parts of that PR still under review. (#5279 — thanks @hartmark)
  • feat(providers): Microsoft 365 Copilot individual provider — adds the copilot-m365-web provider (the 237th), wiring the M365 BizChat framing/connection helpers into a selectable web-session provider backed by m365.cloud.microsoft/chat for individual Microsoft 365 plans. Builds on the M365 pure-framing groundwork from #4696. Regression guard: tests/unit/copilot-m365-web-executor.test.ts. (#5302 — thanks @skyzea1)
🔧 Bug Fixes
  • ci(docker): re-point the Docker Hub / GHCR :latest (and :latest-web) tags to the just-published release. On a release: released event the freshly-created git tag is often not yet visible to git fetch --tags when docker-publish runs, so the :latest-promotion gate built its candidate set purely from git tag -l and resolved the highest semver to the previous version — leaving latest one release behind (3.8.39 published, latest still 3.8.38). The decision now lives in scripts/ci/should-promote-latest.sh, which folds the current VERSION into the candidate set before picking the highest stable semver, making promotion independent of tag-sync timing (a patch published after a higher minor still won't grab latest). Regression guard: tests/unit/build/should-promote-latest-5301.test.ts (#5301)
  • command-code: treat a non-positive max_tokens/max_completion_tokens (e.g. Zoo Code's -1 "let the server choose") as "no limit" — omit the field instead of forcing it to 1. clampMaxTokens previously did Math.max(1, …), so a client -1 was sent upstream as max_tokens: 1, truncating the response to a single token (the observed completion_tokens: 1, content: null, reasoning_content: "The" with finish_reason: stop). Now any value ≤ 0 is dropped so Command Code applies the model's native default; positive values are still floored and clamped to the 200k ceiling. Regression guard: tests/unit/command-code-maxtokens-negative-5166.test.ts (#5166 — thanks @Stazyu)
  • fix(auth): compare-and-swap guard on the OAuth refresh persist — under multi-agent load, the per-connection refresh mutex makes [network refresh + DB write] atomic for one connection, but it does not protect against a third writer (a sibling request, a concurrent HealthCheck, or a replica) landing a fresher refresh_token rotation on the same connection_id between the staleness read and the persist. Overwriting that fresher row reverts the sibling's rotation; the next caller then loads the now-consumed token, Auth0/Anthropic flag it as refresh_token_reused, and the whole token family gets revoked (the 1352× claude/aa5dd5cf invalidation storm). getAccessToken now re-reads the row's current refresh_token immediately before persisting (inside the mutex) and skips the write when it has rotated past the token the caller presented — the caller still receives the freshly-issued access token, only the DB overwrite is skipped. Opt-in via runWithCasGuard (no active guard ⇒ byte-identical behavior); skip/persist counters exposed via getCasGuardStats(). Regression guard: tests/unit/token-refresh-cas-guard-4038.test.ts. (#4038 — thanks @KooshaPari for the root-cause diagnosis)
  • mcp: break the schemas/tools.ts ↔ schemas/toolSearch.ts import cycle introduced when the tool_search defs (#5269) were extracted into their own module — toolSearch.ts imported McpToolDefinition from tools.ts while tools.ts imported toolSearchTool from toolSearch.ts, failing check:cycles on release/v3.8.40. The shared AuditLevel + McpToolDefinition types now live in a leaf schemas/toolDefinition.ts that both import; tools.ts re-exports them for backward compatibility.
  • compression (analytics): record attempted-but-no-op compression runs so Stacked is no longer invisible when it saves nothing. Previously a compression_analytics row was written only on a net-positive saving, so a Stacked (RTK→Caveman) pipeline that ran on already-compact context produced no row — indistinguishable from "never dispatched" (byMode.stacked.count stayed flat while Ultra climbed). Such runs are now recorded with skip_reason and surfaced as a per-mode skipped count plus totalSkipped/bySkipReason in the analytics summary and the Mode Breakdown; the existing net-saving totals/averages are unchanged (skip rows are excluded from them) (#4268 — thanks @abdulkadirozyurt, @androw)
  • cli (tray): fix omniroute server --tray showing no tray on macOS/Linux with no error printed. The wired Unix tray path loaded systray2 through an inline loader that called require("module") inside an ESM .mjs file ("type":"module") → ReferenceError: require is not defined, silently swallowed (regressed in v3.8.34); even if it had loaded, systray2 isn't in node_modules (it's lazily installed into ~/.omniroute/runtime). The loader now delegates to the runtime loader, the icon path (icon.png) is corrected, isTemplateIcon is false (the full-color icon rendered as a white square under macOS template mode), and tray start failures are surfaced to stderr instead of being swallowed (#4605 — thanks @ProgMEM-CC)
  • agent-bridge (antigravity): unwrap the cloudcode-pa .request envelope when converting Antigravity IDE requests. The real IDE sends cloudcode-pa.googleapis.com/v1internal:generateContent with the Gemini request nested under .request ({ project, model, request: { contents, systemInstruction, generationConfig } }), but the bridge read those fields at the top level — yielding an empty conversation, so prompts hung mid-execution. The legacy /v1beta/models/<model>:generateContent top-level shape still works (#4294 — thanks @shabeer)
  • dashboard: add a GitHub releases fallback to the "Update Available" lookup. After the v3.8.28 fix added an npm-registry HTTP fallback, the banner could still stay hidden on networks that reach GitHub (where the news feed already loads) but not registry.npmjs.org. resolveLatestVersion() now tries npm CLI → npm registry → GitHub releases (/repos/diegosouzapw/OmniRoute/releases/latest) before giving up, and logs a warning only when all three fail (#4100)
  • command-code: omit max_tokens when the client omits it so the upstream applies the model's native default, fixing 400 "expected <=200000" on /alpha/generate for high-cap models; an explicit oversized client value is clamped to the 200k endpoint ceiling (#5221 — thanks @adivekar-utexas)
  • combo: wire session stickiness into the round-robin dispatch path. Multi-turn conversations from clients that send no session id (Codex CLI, Claude Code, most OpenAI-compatible tools) were rotated to a different connection on every turn by round-robin combos, busting the upstream prompt-cache → cold high-reasoning starts, intermittent 504s and throughput collapse under concurrency. The weighted/priority paths already honored per-conversation stickiness; the round-robin handler returned before reaching it. Round-robin now starts the rotation at the conversation's sticky connection (failover to the other targets is preserved), and different conversations still spread across connections — only intra-conversation rotation is removed (#5248, #3825 — thanks @bypanghu, @jpsn123, @xz-dev)
  • kiro: replace the synthesized trailing "Continue" turn with a neutral filler ("...") — when an OpenAI→Kiro request ends on an assistant/tool turn, the translator synthesizes the protocol-required trailing user turn, and the literal word "Continue" could be read by Kiro/CodeWhisperer as a real user instruction and trigger unintended agent action. A trailing tool-result turn is still promoted as-is (it already collapses to a real user turn); only the assistant-text-ending case is affected. Regression guards: tests/unit/kiro-continue-filler-5231.test.ts. (#5231)
  • combo: advance to the next combo target on a 400 "requested model is not supported" instead of hard-failing. The 400 guard in the priority strategy treated MODEL_CAPACITY as a block-fallback reason, so a combo that hit a provider lacking a specific model returned a hard 400 even when other targets (different providers) supported it. Such 400s now fall through to the next target. (#5249 — thanks @Chewji9875)
  • dashboard: disabled no-auth providers no longer vanish from the All Providers page. Disabling a no-auth provider (the "No authentication required" toggle, which adds it to blockedProviders) silently removed its card because the page dropped blocked no-auth entries from its render list — the only way back was buried under Settings → Security → Blocked Providers. The page now partitions no-auth entries: visible providers render as before, blocked ones appear in a "Disabled" sub-group with an Enable button that un-blocks them in place. Aggregates, counts and /v1/models still consume the visible-only list (blocked providers stay out of routing). Regression guard: tests/unit/noauth-blocked-partition-5183.test.ts. (#5183, follow-up from #5166 — thanks @WslzGmzs)
  • dashboard: add a parent /dashboard/context page so RSC prefetches of the compression-context hub no longer 404. The route only had sub-routes (settings, combos, ultra, …) and no parent page, so the App Router returned 404 for the bare segment. The parent now redirects to its canonical sub-route (/dashboard/context/settings), honoring a legacy ?tab= query for deep links. Regression guard: tests/unit/dashboard/context-parent-redirect-5298.test.ts (#5298 — thanks @KooshaPari)
  • i18n: add the missing sidebar.gamificationGroup message across all 42 locales — the Gamification sidebar group referenced a titleKey that existed in no locale, logging MISSING_MESSAGE: sidebar.gamificationGroup (en) at runtime (the group still rendered via its titleFallback). The key is now present everywhere so the warning is gone and locale coverage is unaffected (#5298 — thanks @KooshaPari)
  • api(stream): /v1/chat/completions no longer returns SSE for a non-stream OpenAI-compatible request when stream is omitted and the client sends Accept: application/json, text/event-stream — the Vercel AI SDK / OpenAI SDK non-stream signature (doGenerate()/generateText()), which then failed with Invalid JSON response (Unexpected token 'd', "data: {"id"...). The route-level Accept override (#302) and resolveStreamFlag now treat an Accept header that explicitly lists application/json as a JSON opt-in even when it also lists text/event-stream; only a pure Accept: text/event-stream (no application/json) still opts an omitted-stream request into SSE, and an explicit body stream value always wins. The shared decision now lives in acceptHeaderForcesStream. Regression guard: tests/unit/sse-nonstream-accept-5305.test.ts. (#5305 — thanks @md-riaz)
  • providers: drop the retired GPT‑5.2 / GPT‑4.5 models from the direct ChatGPT‑web and Codex surfaces (OpenAI removed them there), so OmniRoute stops advertising/routing models that no longer exist. Scoped on purpose to those two providers — third‑party proxies that still expose the ids are untouched. (#5280 — thanks @backryun)
  • codex: drop the deprecated local_shell hosted tool type before forwarding to OpenAI's Responses API, resolving the omni-combo 400 "The local_shell tool is no longer supported." spike. Inbound Responses local_shell is still accepted and mapped to a caller-side Chat shell function for compatibility. (#5250, #5256 — thanks @KooshaPari)
  • antigravity: retry excluded accounts via the fallback LRU. The combo same-model retry loop accumulates excluded Antigravity connection ids after account-level failures, but auth selection only treated a single excludeConnectionId as a fallback scenario — once exclusions accumulated through excludedConnectionIds, selection could fall back to normal sticky/priority behavior instead of LRU-selecting the next eligible account for the same model/family. Any non-empty accumulated exclude set is now treated as fallback mode. Builds on the family-scoped lockout work in #5180 (v3.8.39). (#5222 — thanks @Ardem2025)
  • grok-cli: strip unsupported sampling params (presencePenalty, frequencyPenalty, logprobs, topLogprobs) before sending to the Grok Build API, fixing 400 'Model does not support parameter presencePenalty' when clients (MiMoCode, Cursor, etc.) send OpenAI-style params. (#5273 — thanks @fulorgnas)
  • grok-cli: accept the full ~/.grok/auth.json object in the dashboard import-token endpoint. The oauthImportTokenSchema only accepted a bare string token while the UI sends the whole auth.json object → 400 Bad Request; the schema now accepts the object and stores the original under providerSpecificData.rawAuthJson for diagnostics and token refresh. (#5258 — thanks @fulorgnas)
  • qoder: coalesce concurrent PAT→job-token exchanges per PAT so high-concurrency / multi-agent bursts no longer stampede openapi.qoder.sh/api/v1/jobToken/exchange before the first exchange populates the completed-token cache; the shared exchange is also decoupled from any single caller's AbortSignal so one aborted waiter can't cancel it for the others. (#5254, #5265 — thanks @KooshaPari)
  • proxy: scope the fallback reachability cache by normalized target URL instead of by hostname, so a failed probe for one endpoint on a shared API host no longer suppresses a later probe for a different endpoint on that same host for the full TTL (host fallback is preserved for malformed URLs). (#5261 — thanks @KooshaPari)
  • proxy: cache failed fast-fail health probes with a short negative TTL instead of the full positive health TTL, so a single transient timeout/load blip no longer marks a working residential SOCKS5 proxy unreachable for the whole window (#5109 regression coverage added). (#5255 — thanks @KooshaPari)
  • mcp: forward HTTP auth to internal tool fetches so MCP tools that call back into the local API surface carry the caller's authorization. (#5218 — thanks @KooshaPari)
  • logging: preserve the outbound provider request headers in the detailed call-log Provider Request payload (previously the upstream response headers were shown there). chatCore now keeps executor-returned request headers when wrapping streaming and non-streaming responses; response headers stay scoped to the Response. (#5257 — thanks @rdself)
  • sse: scope textual <think>/<thinking> tag extraction so generic OpenAI-compatible paths don't rewrite prompt-format content into reasoning_content; an explicit opt-in keeps tag-native families (DeepSeek-R1, QwQ) working while Antigravity/Agy stay excluded by provider/model prefix. (#5224 — thanks @rdself)
  • mcp: break the schemas/tools.ts ↔ schemas/toolSearch.ts import cycle introduced when the tool_search defs (#5269) were extracted into their own module — toolSearch.ts imported McpToolDefinition from tools.ts while tools.ts imported toolSearchTool from toolSearch.ts, failing check:cycles on release/v3.8.40. The shared AuditLevel + McpToolDefinition types now live in a leaf schemas/toolDefinition.ts that both import; tools.ts re-exports them for backward compatibility. (#5282)
  • compression (analytics): record attempted-but-no-op compression runs so Stacked is no longer invisible when it saves nothing. A compression_analytics row was previously written only on a net-positive saving, so a Stacked pipeline that ran on already-compact context produced no row — indistinguishable from "never dispatched". Such runs are now recorded with skip_reason and surfaced as a per-mode skipped count plus totalSkipped/bySkipReason; net-saving totals/averages are unchanged (skip rows excluded). (#5277, #4268 — thanks @abdulkadirozyurt, @androw)
  • cli (tray): fix omniroute server --tray showing no tray on macOS/Linux with no error printed. The Unix tray path loaded systray2 through an inline loader that called require("module") inside an ESM .mjs file → ReferenceError: require is not defined, silently swallowed (regressed in v3.8.34); even if loaded, systray2 is lazily installed into ~/.omniroute/runtime, not node_modules. The loader now delegates to the runtime loader, the icon path is corrected, isTemplateIcon is false (the full-color icon rendered as a white square under macOS template mode), and tray start failures surface to stderr. (#5276, #4605 — thanks @ProgMEM-CC)
  • agent-bridge (antigravity): unwrap the cloudcode-pa .request envelope when converting Antigravity IDE requests. The real IDE sends cloudcode-pa.googleapis.com/v1internal:generateContent with the Gemini request nested under .request, but the bridge read those fields at the top level — yielding an empty conversation, so prompts hung mid-execution. The legacy /v1beta/models/<model>:generateContent top-level shape still works. (#5267, #4294 — thanks @shabeer)
  • dashboard: add a GitHub releases fallback to the "Update Available" lookup. After the v3.8.28 npm-registry fallback, the banner could still stay hidden on networks that reach GitHub but not registry.npmjs.org. resolveLatestVersion() now tries npm CLI → npm registry → GitHub releases before giving up. (#5266, #4100)
  • command-code: omit max_tokens when the client omits it so the upstream applies the model's native default, fixing 400 "expected <=200000" on /alpha/generate for high-cap models; an explicit oversized client value is clamped to the 200k endpoint ceiling. (#5221 — thanks @adivekar-utexas)
🔒 Security
  • authz: require auth for the /v1beta/* Gemini-compatible client API. next.config.mjs rewrote /v1beta/:path*/api/v1beta/:path*, but src/proxy.ts didn't match /v1beta before the rewrite and classifyRoute() didn't classify /api/v1beta/* as client API — so unauthenticated /v1beta/models/...:generateContent traffic could reach the model-serving route without the central client-API auth policy. Both alias and rewritten forms are now classified CLIENT_API, enforcing Bearer auth when REQUIRE_API_KEY is enabled. (#5274 — thanks @rdself)
  • sentinel: security hardening pass across request handling. (#5241 — thanks @iamedwardngo)
  • providers: refresh impersonation User-Agents + TLS fingerprint profiles to current real-client versions; several had drifted or were inconsistent across files, a bot-detection/blocking risk. (#5237 — thanks @backryun)
  • authz (public origin): centralize browser-mutation origin validation into src/server/origin/publicOrigin.ts and wire it through the authz pipeline, replacing the per-route same-origin-only check that 403'd dashboard mutations when served behind a reverse proxy on a different public origin. The module resolves the allowed public origin from configured base-URL env vars or trusted forwarded headers (only when OMNIROUTE_TRUST_PROXY is set and the peer is loopback/LAN via peer-stamp), validates Sec-Fetch-Site metadata, and sanitizes Host/Forwarded inputs (rejects control chars, userinfo, path/query in Host). Regression guards: tests/unit/authz/public-origin.test.ts + tests/unit/authz/pipeline.test.ts. (#5278 — thanks @Thinkscape / @abodera)
📝 Maintenance
  • docs: reorganize docs/, run an accuracy audit, and drop Node 20 from the supported matrix (rebased onto the current release tip). (#5262)
  • providers: remove the discontinued Gemini CLI channel — Google shut it down on 2026-06-18; the supported migration path is Antigravity. (#5246 — thanks @rdself)
  • dashboard: add more provider icons from Lobehub. (#5220 — thanks @backryun)
  • deps: resolve npm install warnings, fix a runtime Cannot find module crash on omniroute serve (the runtime-env module was missing from the npm files allow-list), and clear the 4 moderate npm audit advisories. (#5252 — thanks @yunaamelia), (#5230, #5227)
  • docker: harden the base image against container-scan CVEs and skip comment lines in the #4076 builder heap-ordering check. (#5229, #5233)
  • ci: Trivy advisory scan now ignores unfixed CVEs to cut Security-tab noise. (#5235)
  • test(combo): reconcile the #4279 stop-guard test with the #5249 advance policy. (#5300)
  • test(targets): add 14 unit tests for the shared combo/targetExhaustion.ts handler (provider-exhausted / connection-error / transient rate-limited classification), registered in the vitest discovery list. (#5296 — thanks @KooshaPari)

What's Changed

Full Changelog: https://github.com/diegosouzapw/OmniRoute/compare/v3.8.39...v3.8.40

View originalPermalink
How v3.8.40 went
View all

Discussion