# OmniRoute changelog > Never stop coding. Free MIT AI gateway: one endpoint, 330+ providers (90+ free), 1200+ models — Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 320+ contributors - Vendor: diegosouzapw - Category: AI - Official site: https://omniroute.online/ - Tracked by: What's New (https://whatsnew.fyi/product/omniroute) - Harvested from: GitHub (diegosouzapw/OmniRoute) - Entries below: 10 (newest first) What's New is an index, not a publisher: every entry below links to the vendor's own release notes, which are the authoritative source. Entries are labelled where they are hand-curated sample data, pre-releases, or drawn from a secondary source such as a developer blog. Reuse: the summaries, labels and curation here are © What's New. Quote freely with attribution and a link back; wholesale republication of the corpus is not permitted — terms: https://whatsnew.fyi/terms. The vendors' own release notes remain their publishers'. ## Releases ### v3.8.49 - Date: 2026-07-30 - Version: v3.8.49 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.49 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.49 - **added** — Generalize ensureThinkingBudget to all providers - **added** — Add effort-tier aliases for glm-5.2 and mimo-v2.5 - **added** — Add curated OpenRouter embeddings catalog with specialty merge - **added** — Add opt-in auto-ping to keep Codex quota windows warm - **added** — Add Agnes AI native provider support - **added** — Allow disabling `:` comment heartbeats via SSE configuration - **added** — Add performance.mark/measure to SSE pipeline - **added** — Add Dahl free inference provider - **added** — Add Codex reset credit picker - **added** — Register GPT-5.6 Sol/Terra/Luna model family - **added** — Show Codex plan label in provider and quota views - **added** — Add reorder connections by availability button - **added** — Add 180D and 365D usage and cost analytics periods - **added** — Expose GET /api/usage/model-latency-stats endpoint - **added** — Add compression-mode selector to Context and Cache combos - **added** — Route GitHub Copilot Claude models through native provider - **added** — Add Antigravity reasoning-effort overrides - **added** — Add native xAI Grok Imagine video generation provider - **added** — Add Grok Build CLI tool setup support - **added** — Add Chenzk API OpenAI-compatible gateway > **All 1383 entries from this cycle are listed below**, one line each — descriptions are > trimmed to fit GitHub's 125,000-character release body. Full wording, context and links: > [CHANGELOG.md](https://github.com/diegosouzapw/OmniRoute/blob/main/CHANGELOG.md#3849--2026-07-28). _Living section — regenerated 2026-07-19 from all 306 cycle commits (bump 2c62333b0 → tip). Bullets carry the merged PR and its author; direct pushes listed separately. Finalized at the v3.8.49 release._ ###### ✨ New Features - **feat:** generalize ensureThinkingBudget to all providers +… (#6979) — @rafaumeu - **feat(6922):** effort-tier aliases for glm-5.2 & mimo-v2.5 on… (#6987) — @rafaumeu - **feat(providers):** curated OpenRouter embeddings catalog + specialty merge… (#6994) - **feat(quota):** opt-in auto-ping to keep Codex quota windows warm (#6995) - **feat(providers):** add Agnes AI native provider support (#7035) — @HouMinXi - **feat(sse):** allow disabling `:` comment heartbeats via… (#7036) — @xier2012 - **feat(perf):** add performance.mark/measure to SSE pipeline +… (#7045) — @oyi77 - **feat(providers):** add Dahl free inference provider (#7062) — @growab - **feat(ci):** boot-smoke the packed npm tarball (check:pack-boot,… (#7086) - **feat(ci):** hotfix fast-lane + tests-only E2E skip (WS3.1) (#7088) - **feat(ci):** continuous release-green — on-push quick gate + 3x/day… (#7089) - **feat(ci):** duration-balanced E2E shards via LPT bin-packing (WS4.1) (#7090) - **feat(ci):** TypeScript 7 native shadow for typecheck:core (WS4.2,… (#7091) - **feat(release):** npm staged publishing + pre-publish boot-smoke (WS1.3) (#7092) - **feat(release):** post-publish verifier — clean-container install + boot… (#7109) - **feat(ci):** Mergify merge queue + manual-train fallback runbook… (#7112) - **feat(ci):** Windows leg for Electron prepare smoke (WS1.5) (#7113) - **feat(ci):** Codecov patch coverage (informational) + fix missing… (#7114) - **feat(sidecar):** support conditional provider manifest refresh (#7130) — @KooshaPari - **feat(homolog):** real-environment E2E homologation suite (npm run… (#7133) - **feat(usage):** add Codex reset credit picker (#7154) — @JxnLexn - **feat(ci):** Trunk Flaky Tests uploads for vitest + Playwright E2E… (#7175) - **feat(ci):** Trunk Flaky Tests upload on the fast-path vitest job… (#7205) - **feat(kiro):** register GPT-5.6 Sol/Terra/Luna model family (#7209) - **feat(dashboard):** show Codex plan label in provider and quota views (#7210) - **feat(dashboard):** add reorder connections by availability button (#7211) - **feat(dashboard):** add 180D and 365D usage/cost analytics periods (#7213) - **feat(api):** add Vary: Accept-Encoding to token-authenticated /v1*… (#7217) - **feat(api):** expose GET /api/usage/model-latency-stats (#7218) - **feat(dashboard):** add compression-mode selector to Context & Cache combos… (#7219) - **feat(sse):** route GitHub Copilot Claude models through native… (#7223) - **feat(mitm):** add Antigravity reasoning-effort overrides (#7228) - **feat:** replace free-text model inputs with hidePaid-aware… (#7229) - **feat:** editable ComfyUI base-URL field + per-connection… (#7232) - **feat(sse):** add optional-enum null-omission idiom for codex… (#7233) - **feat(sse):** preserve tools/tool_choice for tool-bearing requests… (#7235) - **feat(api):** accept x-goog-api-key header for client-facing auth (#7236) - **feat(sse):** add native xAI Grok Imagine video generation provider (#7238) - **feat:** add Type filter and easiest-first sort to Free Provider… (#7240) - **feat(cli):** add Grok Build CLI tool setup (~/.grok/config.toml) (#7241) - **feat(provider):** add Chenzk API OpenAI-compatible gateway (#7246) - **feat(providers):** let custom connections opt into prompt-cache capability (#7257) - **feat(db):** include xp_audit_log in automatic retention/prune (#7260) - **feat(api):** structured X-Routing-Fallback-Reason header for relay… (#7262) - **feat(compression):** support RTK TOML schema v1 filt _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.49]_ ### v3.8.48 - Date: 2026-07-13 - Version: v3.8.48 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.48 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.48 - **fixed** — Ship dist/head-response-guard.cjs in the npm tarball to prevent ERR_MODULE_NOT_FOUND crashes on every omniroute boot - **fixed** — Fix Electron Windows packaging by spawning npx.cmd through a shell for better-sqlite3 ABI rebuilds - **fixed** — Fix Sonar quality gate by ensuring coverage lcov reaches the scanner at coverage/lcov.info - **fixed** — Await async isCloudEnabled() gate in Kiro auto-import route so cloud sync respects disabled state - **fixed** — Replace dead structuredClone fallback with real JSON fallback in reasoning-split clone - **fixed** — Handle async reader.cancel() rejection in codex executor - **fixed** — Use deterministic localeCompare for sorting - **fixed** — Add path-traversal guard in classify-pr-changes.mjs - **fixed** — Use npm's bundled node-gyp instead of npx --yes for Docker better-sqlite3 rebuild - **changed** — Codex bulk-import endpoint now accepts 9router's camelCase account export (accessToken/refreshToken/idToken/expiresAt) alongside snake_case format - **added** — Langfuse observability plugin - **added** — Context requirements config for per-target filtering in combos - **added** — Icons for 46 providers that were missing images - **added** — Vendored GCF (Headroom) codec updated to spec v3.2 with nested flattening - **added** — Shorthand proxy formats and protocol header mode for bulk import - **added** — OpenVecta AI inference gateway provider - **added** — Traditional Chinese (zh-TW) localization for frontend and CLI - **added** — Route xAI clients to Grok's native /v1/responses endpoint - **added** — Per-model web-search and web-fetch interception rules - **added** — Sidebar quick-filter search input to filter nav sections and items client-side by label > ⚠️ **Hotfix release.** The published npm package for 3.8.47 crashed on every boot ([#7065](https://github.com/diegosouzapw/OmniRoute/issues/7065)) and was deprecated — **3.8.48 is the first installable release of the v3.8.47 cycle**, so everything listed under [3.8.47] below ships here. ###### 🐛 Bug Fixes - **fix(build):** ship `dist/head-response-guard.cjs` in the npm tarball — the prepublish prune allowlist lacked it, so every `omniroute` boot of the published 3.8.47 crashed with `ERR_MODULE_NOT_FOUND` (3rd occurrence of this class after tls-options/3.8.41); now allowlisted, enforced by `check:pack-artifact`, and guarded by a closure test that derives every `server-ws.mjs` sibling import ([#7065](https://github.com/diegosouzapw/OmniRoute/issues/7065), [#7040](https://github.com/diegosouzapw/OmniRoute/issues/7040)) - **fix(build):** Electron Windows packaging — the better-sqlite3 Electron-ABI rebuild now spawns `npx.cmd` through a shell (Node's CVE-2024-27980 hardening made the shell-less spawn fail with `status null` on Windows runners, breaking the v3.8.47 desktop build) - **fix(ci):** Sonar quality gate zeroed on new code — the coverage lcov now reaches the scanner at `coverage/lcov.info` (it read 0% on every scan), the async `isCloudEnabled()` gate in the Kiro auto-import route is awaited (cloud sync ran even when disabled), the dead `structuredClone` fallback in the reasoning-split clone is a real JSON fallback, the codex executor handles the async `reader.cancel()` rejection, deterministic `localeCompare` sorts, a path-traversal guard in `classify-pr-changes.mjs`, and the Docker better-sqlite3 rebuild uses npm's bundled node-gyp instead of `npx --yes` - **chore(ci):** the Sonar quality gate is informational (`sonar.qualitygate.wait=false`) while the org's SonarCloud plan cannot associate the tuned "OmniRoute way" gate (coverage ≥60 aligned with the repo floor) --- #### 📦 Everything from the v3.8.47 cycle ships here _The 3.8.47 npm package was never installable (#7065), so **3.8.48 is the release that actually delivers the whole v3.8.47 cycle** — full notes below:_ - **9router Codex import**: the Codex bulk-import endpoint (`POST /api/oauth/codex/import`) now accepts 9router's camelCase account export (`accessToken`/`refreshToken`/`idToken`/`expiresAt` + nested `providerSpecificData`), not just snake_case — `normalizeCodexImportRecord` maps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged ([#6665](https://github.com/diegosouzapw/OmniRoute/issues/6665)) — thanks @deadcoder0904. Regression guard: `tests/unit/codexBulkImport.test.ts` (9router camelCase record, pre-supplied `providerSpecificData` without an id_token, snake_case-not-overridden, and a full `{accounts:[...]}` flatten). ###### ✨ New Features - **feat(plugins):** Langfuse observability plugin. ([#6577](https://github.com/diegosouzapw/OmniRoute/pull/6577) — thanks @chirag127) - **feat(combo):** context requirements config for per-target filtering in combos. ([#6907](https://github.com/diegosouzapw/OmniRoute/pull/6907) — thanks @oyi77) - **feat(providers):** icons for 46 providers that were missing images. ([#6926](https://github.com/diegosouzapw/OmniRoute/pull/6926) — thanks @oyi77) - **feat(compression):** vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). ([#6838](https://github.com/diegosouzapw/OmniRoute/pull/6838) — thanks @blackwell-systems) - **feat(proxy):** shorthand proxy formats + protocol header mode for bulk import. ([#6867](https://github.com/diegosouzapw/OmniRoute/pull/6867) — thanks @growab) - **feat(provider):** OpenVecta AI inference gateway. ([#6833](https://github.com/diegosouzapw/OmniRoute/pull/6833) — thanks @hajilok) - **feat(i18n):** Traditional Chinese (zh-TW) localization for frontend and CLI. ([#6320](https://github.com/diegosouzapw/OmniRoute/pull/6320) — thanks @lunkerchen) - **feat(xai):** route xAI clients to Gr _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.48]_ ### v3.8.47 - Date: 2026-07-13 - Version: v3.8.47 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.47 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.47 - **added** — Langfuse observability plugin - **added** — Context requirements config for per-target filtering in combos - **added** — Icons for 46 providers that were missing images - **changed** — Vendored GCF (Headroom) codec updated to spec v3.2 with nested flattening - **added** — Shorthand proxy formats and protocol header mode for bulk import - **added** — OpenVecta AI inference gateway provider - **added** — Traditional Chinese (zh-TW) localization for frontend and CLI - **added** — Route xAI clients to Grok's native /v1/responses endpoint - **added** — Per-model web-search and web-fetch interception rules - **added** — Changelog fragments system in changelog.d/ to eliminate CHANGELOG merge conflicts - **added** — validate-release-green --full-ci command to reproduce entire ci.yml static gate set locally - **added** — Sidebar quick-filter search input in dashboard that filters nav sections client-side by label - **added** — Auto combos gain strict budget-cap fallback policy with X-OmniRoute-Budget-Fallback header - **added** — Config-driven parameter denylist and allowlist per provider and model with auto-learn from upstream 400s - **added** — Per-combo reasoning token buffer toggle checkbox in combo builder - **added** — Routing Strategy settings card on Settings page plus per-provider account-routing override - **added** — Combo-level sticky round-robin limit setting resolved via cascade from per-combo to global combo sticky to account sticky - **changed** — Codex bulk-import endpoint now accepts 9router's camelCase account export format alongside snake_case - **9router Codex import**: the Codex bulk-import endpoint (`POST /api/oauth/codex/import`) now accepts 9router's camelCase account export (`accessToken`/`refreshToken`/`idToken`/`expiresAt` + nested `providerSpecificData`), not just snake_case — `normalizeCodexImportRecord` maps the camelCase aliases onto the existing snake_case keys, filling each only when absent so snake_case/mixed exports keep working unchanged ([#6665](https://github.com/diegosouzapw/OmniRoute/issues/6665)) — thanks @deadcoder0904. Regression guard: `tests/unit/codexBulkImport.test.ts` (9router camelCase record, pre-supplied `providerSpecificData` without an id_token, snake_case-not-overridden, and a full `{accounts:[...]}` flatten). ###### ✨ New Features - **feat(plugins):** Langfuse observability plugin. ([#6577](https://github.com/diegosouzapw/OmniRoute/pull/6577) — thanks @chirag127) - **feat(combo):** context requirements config for per-target filtering in combos. ([#6907](https://github.com/diegosouzapw/OmniRoute/pull/6907) — thanks @oyi77) - **feat(providers):** icons for 46 providers that were missing images. ([#6926](https://github.com/diegosouzapw/OmniRoute/pull/6926) — thanks @oyi77) - **feat(compression):** vendored GCF (Headroom) codec updated to spec v3.2 (nested flattening). ([#6838](https://github.com/diegosouzapw/OmniRoute/pull/6838) — thanks @blackwell-systems) - **feat(proxy):** shorthand proxy formats + protocol header mode for bulk import. ([#6867](https://github.com/diegosouzapw/OmniRoute/pull/6867) — thanks @growab) - **feat(provider):** OpenVecta AI inference gateway. ([#6833](https://github.com/diegosouzapw/OmniRoute/pull/6833) — thanks @hajilok) - **feat(i18n):** Traditional Chinese (zh-TW) localization for frontend and CLI. ([#6320](https://github.com/diegosouzapw/OmniRoute/pull/6320) — thanks @lunkerchen) - **feat(xai):** route xAI clients to Grok's native `/v1/responses` endpoint. ([#6709](https://github.com/diegosouzapw/OmniRoute/pull/6709) — thanks @diegosouzapw) - **feat(routing):** per-model web-search/web-fetch interception rules. ([#3384](https://github.com/diegosouzapw/OmniRoute/issues/3384), [#6814](https://github.com/diegosouzapw/OmniRoute/pull/6814) — thanks @diegosouzapw) - **feat(release):** `changelog.d/` fragments — eliminates the CHANGELOG merge-storm cascade. ([#6783](https://github.com/diegosouzapw/OmniRoute/pull/6783) — thanks @diegosouzapw) - **feat(quality):** `validate-release-green --full-ci` reproduces the entire ci.yml static gate set locally. ([#6583](https://github.com/diegosouzapw/OmniRoute/pull/6583) — thanks @diegosouzapw) - **feat(dashboard):** sidebar quick-filter — a search input at the top of the expanded dashboard sidebar (`src/shared/components/Sidebar.tsx`) filters nav sections/groups/items client-side by label as you type, reusing the existing `common.search`/`common.noResults` i18n keys (zero new locale edits) and the shared `Input` `icon="search"` pattern; matching sections auto-expand while searching (bypassing the accordion/pin state) and collapse back to normal once the query is cleared. Pure filtering logic extracted into `filterSidebarSectionsByQuery()` (`src/shared/utils/sidebarSearch.ts`) for isolated unit testing. Regression guard: `tests/unit/sidebar-search-filter.test.ts`, `src/shared/components/Sidebar.search.test.tsx`. (#4013 — thanks @crochabe-cyber) - **feat(combo):** `auto/*` combos gain a strict budget-cap fallback policy — `X-OmniRoute-Budget-Fallback: strict` (or the persisted `config.budgetFallback: "strict"`) makes an over-budget request fail fast with `HTTP 402` instead of the previous silent fallback to the globally cheapest candidate, which could still exceed the cap. The default (`cheapest`) preserves existing behavior. Builds on the existing `X-OmniRoute-Budget`/`X-OmniRoute-Mode` per-request controls (#6023/#6024/#6025), consolidated into `resolveRequestAutoControls()`. Regression guard: `tests/unit/auto-combo-budget-fallback-3470.test.ts`. (#3470) - **Provid _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.47]_ ### v3.8.46 - Date: 2026-07-07 - Version: v3.8.46 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.46 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.46 - **added** — Hide paid-only models from auto/* routing when hidePaidModels setting is enabled - **added** — Add provider-family auto combos (auto/glm, auto/minimax, auto/mimo, auto/zai, auto/gemma, auto/llama, auto/gemini) that materialize virtual combos spanning installed backends for each model family - **added** — Implement native proxy-pool round-robin and egress IP rotation with multiple proxies per scope and rotation strategies (round-robin, random, sticky-per-N-min) - **added** — Add end-to-end tool and function calling support on the native Gemini /v1beta endpoint in both request and response directions - **added** — Add enterprise/work tier support for copilot-m365-web provider via M365_ENTERPRISE_OVERRIDES preset and agent field configuration ###### ✨ New Features - **feat(sse):** **hide paid-only models from `auto/*` routing** when `hidePaidModels` is on ([#6512](https://github.com/diegosouzapw/OmniRoute/issues/6512)) — follow-up to #6328/#6495. PR #6495 hid paid-only models from the `GET /v1/models` listing, but `auto/*` combos (`auto/best-coding`, `auto/glm`, …) could still pick a paid-only backend into their candidate pool → a 402/403 at request time. `createVirtualAutoCombo` now filters the candidate pool through the new pure `open-sse/services/autoCombo/paidModelFilter.ts` (`filterPaidOnlyCandidates`), applying the same free-model predicate #6495 uses in `catalog.ts` (`providerHasFreeModels(provider) && isFreeModel(provider, {id})`) whenever `settings.hidePaidModels === true`. Applied before the category/tier/family narrowing, so it covers every `auto/*` combo; an all-paid pool degrades to the existing graceful empty-pool path. **Opt-in — default OFF leaves the pool unchanged** (identity). Regression guard: `tests/unit/autoCombo/paid-model-filter-6512.test.ts` (4, incl. the default-off identity guard). - **feat(sse):** **provider-family auto combos** — `auto/glm`, `auto/minimax`, `auto/mimo`, `auto/zai`, `auto/gemma`, `auto/llama`, `auto/gemini` ([#6453](https://github.com/diegosouzapw/OmniRoute/issues/6453)) — new routable ids that materialize an on-demand virtual combo spanning whatever installed backends currently expose that model family, degrading gracefully as backends rotate. A new pure `open-sse/services/autoCombo/modelFamily.ts` (`detectModelFamily`) classifies by model-id prefix for six families; `zai` is instead resolved by provider id (z.ai's hosted API serves the same `glm-*` model ids as every other GLM backend, so `auto/zai` means "route to my z.ai backend specifically" vs `auto/glm`'s "any connected GLM backend"). Reuses the existing `createVirtualAutoCombo` on-demand materialization path (no DB writes) and the `/v1/models` catalog advertising loop. Regression guard: `tests/unit/autoCombo/provider-family-combos.test.ts` (11). - **feat(proxy):** native **proxy-pool round-robin / egress IP rotation** ([#6365](https://github.com/diegosouzapw/OmniRoute/issues/6365)) — a scope (global / provider / account) can now hold **multiple** proxies as a pool with a rotation strategy, so outbound requests cycle their egress IP instead of pinning one proxy per scope. Migration `117_proxy_pool_rotation.sql` lifts the `UNIQUE(scope, scope_id)` constraint (rebuild via the canonical rename/copy/drop; existing single assignments become 1-element pools) and adds a `proxy_scope_rotation` companion table holding the per-scope strategy + a persisted monotonic round-robin cursor. Strategies: `round-robin` (default, monotonic cursor — never `Math.random`), `random`, and `sticky-per-N-min`. Resolution (`resolveProxyForScopeFromRegistry` / `resolveProxyForConnectionFromRegistry`) now fetches the alive, position-ordered candidate set (unchanged `PROXY_ALIVE_PREDICATE`) and applies the strategy; an empty / all-dead pool still returns `null` — the #6246 fail-closed guard is **untouched** (never falls through to direct egress). Backend + DB only; dashboard pool-builder UI is a follow-up. Regression guard: `tests/unit/proxy-pool-rotation-6365.test.ts` (8, incl. fail-closed + backward-compat). - **feat(providers):** end-to-end **tool/function calling on the native Gemini `/v1beta` endpoint** ([#6222](https://github.com/diegosouzapw/OmniRoute/issues/6222)) — both directions of the Gemini↔OpenAI conversion now preserve tool data (previously silently dropped). Request side: `convertGeminiToInternal` (extracted to its own testable module) maps `tools[].functionDeclarations` → OpenAI `tools`, prior `functionCall` parts → assistant `tool_calls`, and `functionResponse` parts → `tool`-role messages. Response side: `convertOpenAIResponseToGemini` emits `parts[].functionCall {name,args}` from `message.tool_calls`, and the streaming `openAIChunkToGeminiChunk` accumulates fragmented `tool _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.46]_ ### v3.8.45 - Date: 2026-07-06 - Version: v3.8.45 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.45 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.45 - **added** — Add Yuanbao (web) as a cookie-session provider with support for Tencent Yuanbao, DeepSeek V3/R1, and Hunyuan models - **added** — Route the built-in agentrouter through the dynamic Claude-Code wire image while preserving its own registry baseUrl and x-api-key auth - **added** — Enable bulk-add API keys for Cloudflare Workers AI with per-entry providerSpecificData to avoid shared-object reuse - **added** — Show effective routing share percentage next to each weight in weighted combos when weights don't sum to 100 - **changed** — Rename the status widget's 'Cloud Sync' label to 'Remote Settings Sync' - **added** — Add opt-in advanced base-URL override for built-in providers hidden behind an Advanced toggle - **added** — Add an option to disable session stickiness per-combo or globally to allow round-robin or random combos to rotate connections on every request - **added** — Add the OMNIROUTE_NO_SUDO env flag for root-less or user-namespaced deployments - **added** — Add Requesty as an OpenAI-compatible gateway provider with ~200 free requests per day - **added** — Add configured-only and available-only filters to the Free Provider Rankings page via server-side query params ###### ✨ New Features - **feat(providers):** add **Yuanbao (web)** as a cookie-session provider ([#6196](https://github.com/diegosouzapw/OmniRoute/issues/6196)) — `yuanbao-web` (Tencent Yuanbao, `yuanbao.tencent.com`) with cookie-only auth (`hy_user`/`hy_token` + public agent id), SSE→OpenAI translation incl. `reasoning_content`, exposing DeepSeek V3/R1 + Hunyuan / Hunyuan-T1. Regression guard: `tests/unit/providers-yuanbao-web.test.ts`. `together-web` was **deferred** (no verifiable web-session endpoint — needs a captured request) and `huggingchat-web` **dropped** (the existing `huggingchat` already is a web-cookie provider). (thanks @chirag127) - **feat(providers):** route the built-in **agentrouter** through the dynamic Claude-Code wire image ([#6056](https://github.com/diegosouzapw/OmniRoute/issues/6056)) — a small static allow-set (`CC_WIRE_IMAGE_BUILTINS` in `open-sse/services/ccWireImageBuiltins.ts`), consulted by `isClaudeCodeCompatible` / `isClaudeCodeCompatibleProvider` / `applyFingerprint`, makes agentrouter adopt the CC wire-image headers + fingerprint **while guarding the CC baseUrl/auth branches** so it keeps its own registry `baseUrl` and `x-api-key` auth. Regression guard: `tests/unit/agentrouter-cc-wire-image.test.ts` (asserts the wire image is applied AND agentrouter's baseUrl/auth are preserved). Live WAF-acceptance against agentrouter.org is a VPS validation follow-up (Hard Rule #18). - **feat(providers):** **bulk-add API keys for Cloudflare Workers AI** ([#6174](https://github.com/diegosouzapw/OmniRoute/issues/6174)) — `cloudflare-ai` is removed from the bulk-add exclusion list and the bulk parser gains a 3-field `name|accountId|apiKey` mode; the bulk route now builds a **per-entry** `providerSpecificData` so each key carries its own `accountId` (fixing the previous shared-object reuse), and both the create + key-validation paths receive it. Regression guard: `tests/unit/bulk-api-key-parser-cloudflare.test.ts`. (thanks @muflifadla38) - **feat(dashboard):** routing/settings UX clarity ([#6147](https://github.com/diegosouzapw/OmniRoute/issues/6147)) — (1) weighted combos show the **effective routing share %** next to each weight when weights don't sum to 100 (`WeightTotalBar.tsx`); (2) the status widget's user-facing **"Cloud Sync" label is renamed** to "Remote Settings Sync" (`CloudSyncStatus.tsx`; internal ids/state untouched); (3) built-in providers gain an **opt-in advanced base-URL override** (`isBaseUrlOverrideEligibleProvider`, hidden behind an "Advanced" toggle, reusing the existing `providerSpecificData.baseUrl` persistence — not globally widened). Regression guard: `tests/unit/routing-settings-ux-6147.test.ts`. - **feat(combo):** add an option to **disable session stickiness**, per-combo or globally — round-robin / random combos can rotate to a different connection on every request instead of pinning a whole conversation to one connection by its first-message hash. Resolution precedence per-combo `config.disableSessionStickiness` → global `settings.disableSessionStickiness` → default `false` (preserves the #3825 prompt-cache/504 fix); gates **both** stickiness call sites in `open-sse/services/combo.ts`. Exposed as a global toggle (Combo Defaults) and a per-combo Inherit/on/off control. ([#6168](https://github.com/diegosouzapw/OmniRoute/issues/6168)) Regression guard: `tests/unit/combo-disable-session-stickiness.test.ts`. (thanks @RCrushMe) - **feat(docker):** add the `OMNIROUTE_NO_SUDO` env flag for root-less / user-namespaced deployments — the MITM cert-trust command path (`resolveSudoSpawn` in `src/mitm/systemCommands.ts`) now strips the leading `sudo` when the flag is truthy, in addition to the existing root / sudo-missing cases, so the Proxy Agent runs without `sudo` (the operator trusts the CA manually, e.g. via `NODE_EXTRA_CA_CERTS`). Argv-array `spawn` preserved — no shell interpolation (Hard Rule #13). ([#6122](https://github.com/diegosouzapw/OmniRoute/issues/6122)) Regression guard: `test _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.45]_ ### v3.8.44 - Date: 2026-07-04 - Version: v3.8.44 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.44 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.44 - **added** — Add throttling for upstream quota fetches on the per-request preflight path via a global min-interval gate, configurable via OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS (default 250ms, clamped 0..5000) - **added** — Add per-request Auto-Combo controls via X-OmniRoute-Mode and X-OmniRoute-Budget headers to steer scoring and set a hard per-request USD cost ceiling - **added** — Add the Kenari OpenAI-compatible gateway (BYOK) as a provider - **added** — Add claude-sonnet-5 to the Antigravity model catalog - **added** — Add /v1/ocr endpoint (Mistral OCR), an OCR provider category, and Mistral moderation support - **added** — Add discoveryResults DB module with CRUD operations and persist provider-discovery findings through the discovery_results table - **added** — Add /api/discovery/* HTTP surface under strict loopback-only authorization for running provider scans and managing discovery results - **added** — Add a dashboard UI tab (Tools → Discovery) to run scans and review, verify, or delete provider discovery findings - **added** — Expose a read-only provider plugin manifest at GET /api/v1/provider-plugin-manifest for sidecar/relay discovery - **added** — Advertise the provider manifest URL to Bifrost/CLIProxyAPI via the X-OmniRoute-Provider-Manifest-Url header - **added** — Add a latency/speed-optimized routing mode with rankBySpeed scoring core and the omniroute_pick_fastest_model MCP tool - **changed** — Refresh The Old LLM (Free) model catalog with current free tier models (GPT-5/5.1/5.2/5.3/5.4, o3/o4-mini, Gemini 3 Pro / 2.5 Pro / 2.0 Flash / 1.5 Flash, Claude 4.6 Opus/Sonnet & 4.5 Haiku, GPT-4o, Grok 4, DeepSeek V3/R1, Sonar Pro) while keeping legacy alias IDs - **fixed** — Fix mapModel() to pass known upstream IDs through unchanged so Gemini/o-series/Grok/DeepSeek/Sonar models no longer collapse onto GPT_5_4 - **added** — Surface Codex banked reset credits per connected account by reading rate_limit_reset_credits.available_count from the /wham/usage payload and rendering a Banked Reset Credits row on the provider-limits dashboard ###### ✨ New Features - **feat(resilience):** throttle upstream quota fetches on the per-request preflight path ([#6009](https://github.com/diegosouzapw/OmniRoute/issues/6009)) — a new global min-interval gate (`open-sse/services/quotaFetchThrottle.ts`) spaces the actual network calls made by the Codex quota fetcher so that many accounts on one IP no longer fetch quota in the same second (which, per `router-for-me/CLIProxyAPI#2385`, can get a Codex OAuth token revoked). Complements the existing bulk-sync spacing (`PROVIDER_LIMITS_SYNC_SPACING_MS`) which already serialized the periodic provider-limits sync — this covers the concurrent combo/preflight path it didn't. Cache hits are never delayed; fail-open (only ever awaits a timer). Configurable via `OMNIROUTE_QUOTA_FETCH_MIN_INTERVAL_MS` (default 250ms, clamped 0..5000; `0` disables). Regression guard: `tests/unit/quota-fetch-throttle-6009.test.ts` (5). (thanks @powellnorma) - **feat(autoCombo):** add **per-request Auto-Combo controls** via two headers ([#6024](https://github.com/diegosouzapw/OmniRoute/issues/6024) / [#6025](https://github.com/diegosouzapw/OmniRoute/issues/6025) / [#6023](https://github.com/diegosouzapw/OmniRoute/issues/6023)) — `X-OmniRoute-Mode` steers an `auto` combo's scoring for a single request (friendly presets `fast`/`balanced`/`quality`/`cheap`/`reliable`/`offline` **or** a raw mode-pack name; `balanced` forces the default weights), and `X-OmniRoute-Budget` sets a hard per-request USD cost ceiling. Both override the combo's stored config only for the request that carries them; unknown/garbage values are ignored so the saved config is preserved. The resolvers are pure (`open-sse/services/autoCombo/requestControls.ts`) and feed the engine's existing `config.modePack` / `config.budgetCap` inputs — no engine changes. Regression guard: `tests/unit/auto-combo-request-controls-6024.test.ts` (5). (thanks @chirag127) - **feat(providers):** add the **Kenari** OpenAI-compatible gateway (BYOK). Regression guard: `tests/unit/kenari.test.ts`. (thanks @doedja) - **feat(models):** add `claude-sonnet-5` to the Antigravity model catalog (alias mapping in `antigravityModelAliases.ts`) ([#6103](https://github.com/diegosouzapw/OmniRoute/pull/6103)). Regression guard: `tests/unit/antigravity-model-aliases.test.ts`. (thanks @anki1kr) - **feat(api):** add `/v1/ocr` endpoint (Mistral OCR), an OCR provider category, and Mistral moderation support. ([#5950](https://github.com/diegosouzapw/OmniRoute/pull/5950)) (thanks @waguriagentic) - **Discovery tool (Phase 2):** add the `discoveryResults` DB module (CRUD over the `discovery_results` table, migration 074) and wire the opt-in provider-discovery service to persist and read findings through it (`persistDiscoveryResult`, `getDiscoveryResults`, `getDiscoveryResultById`, `markVerified`, `deleteDiscoveryResult`) with `(provider, method, endpoint)` upsert de-duplication. Adds the `/api/discovery/*` HTTP surface — `GET /results`, `GET|DELETE /results/:id`, `POST /scan`, `POST /verify/:id` — under **strict loopback-only** authorization (`/api/discovery/` is in `LOCAL_ONLY_API_PREFIXES` and is NOT manage-scope-bypassable, so the `scan` route's outbound probes can never be reached from a tunnel/remote origin). Adds a **dashboard UI tab** (Tools → Discovery, `/dashboard/discovery`) to run scans and review, verify, or delete findings. The service stays **opt-in / default-off**. ([#5939](https://github.com/diegosouzapw/OmniRoute/pull/5939)) - **feat(api):** expose a read-only provider plugin manifest at `GET /api/v1/provider-plugin-manifest` for sidecar/relay discovery. ([#6001](https://github.com/diegosouzapw/OmniRoute/pull/6001)) (thanks @KooshaPari) - **feat(sidecar):** advertise the provider manifest URL to Bifrost/CLIProxyAPI via the `X-OmniRoute-Provider-Manifest-Url` header (`OMNIROUTE_PROVIDER_MANIFEST_URL`). ([#6007](https://github.com/diegosouzapw/OmniRoute/pull/6007)) (thanks @KooshaPari) - **feat(autoCombo):** add a latency/spe _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.44]_ ### v3.8.43 - Date: 2026-07-02 - Version: v3.8.43 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.43 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.43 - **added** — Usage endpoint and @@om-usage command now report personal API-key quotas as remaining percentages with provider USD cost drilldown via /api/usage/provider-window-costs endpoint - **added** — Usage quota system detects observed provider quota resets by same resetAt values instead of relying only on recorded weekly events - **added** — Live dashboard WebSocket can be fronted by reverse proxy or Cloudflare Tunnel via NEXT_PUBLIC_LIVE_WS_PUBLIC_URL environment variable with runtime support for prebuilt Docker and npm images - **added** — Handshake endpoint /api/v1/ws?handshake=1 now echoes lazily-read live.publicUrl for runtime WebSocket URL resolution - **added** — Optional auto-sync feature for CLI tool profiles after provider model sync, supporting Codex and Claude Code profiles via OMNIROUTE_AUTO_SYNC_CODEX_PROFILES and OMNIROUTE_AUTO_SYNC_CLAUDE_PROFILES feature flags - **added** — CLI Code dashboard now includes CLI profile auto-sync card to toggle Codex and Claude profile auto-synchronization - **added** — serve command startup banner now displays running OmniRoute version beneath ASCII logo - **added** — Analytics dashboard shows $0 cost for flat-rate subscription providers instead of inflated per-token estimates, covering ChatGPT Web, grok-web, Minimax Coding, Kimi Coding, GLM Coding, Alibaba Coding Plan, and Xiaomi MiMo - **added** — New MCP tools omniroute_rtk_discover and omniroute_rtk_learn expose RTK tool-output learn and discover workflow for agents to grow RTK filter catalog - **changed** — Provider quota remaining is scaled by configured quota cutoff so protected reserve reads as 0% left - **changed** — Analytics surfaces now support opt-in flatRateAsZero cost option while budget, quota, and routing continue estimating unchanged for flat-rate providers ##### [3.8.43] — 2026-07-02 ###### ✨ New Features - **usage (quota percentages + provider USD drilldown):** `@@om-usage` and the HTTP usage endpoint now report personal API-key quotas as **remaining percentages** (USD amounts stay out of the command output), provider quota remaining is scaled by the configured quota cutoff so the protected reserve reads as 0% left, and the quota dashboard regains a **provider USD cost drilldown** (`/api/usage/provider-window-costs` + `ProviderUsdCostModal`, management-auth gated). Also honors **observed provider quota resets**: a same-`resetAt` reset (usage dropping back to the reset floor) is detected and preferred over stale recorded weekly events for provider USD windows and API-key USD quotas. New `src/lib/usage/providerWindowCosts.ts`. Regression guards: `tests/unit/provider-window-costs.test.ts`, `tests/unit/internal-usage-command.test.ts`, `tests/unit/api-key-usage-limits.test.ts`, `tests/unit/lib/quota-reset-events.test.ts`. Extracted from [#5863](https://github.com/diegosouzapw/OmniRoute/pull/5863) by [@Witroch4](https://github.com/Witroch4). - **dashboard (live WS behind reverse proxy):** the live dashboard WebSocket can now be fronted by a reverse proxy or Cloudflare Tunnel via `NEXT_PUBLIC_LIVE_WS_PUBLIC_URL` (e.g. `wss://ws.my-ai.com/live-ws`). The URL is honored both at build time (env inlined into the bundle) and at **runtime** for prebuilt Docker/npm images: the `/api/v1/ws?handshake=1` handshake now echoes a lazily-read `live.publicUrl` (only `ws://`/`wss://` values are accepted; anything else is rejected to `null`), and `useLiveDashboard` resolves the URL from that handshake before connecting, falling back to the previous `ws(s)://hostname:20129` default. Also documents `LIVE_WS_ALLOWED_HOSTS` and aligns the GitLab Duo OAuth scopes line in `.env.example` with the live config (`ai_features read_user`). Regression guard: `tests/unit/live-ws-public-url.test.ts` (5). ([#5877](https://github.com/diegosouzapw/OmniRoute/pull/5877) by [@ianriizky](https://github.com/ianriizky)) - **providers (CLI profile auto-sync):** opt-in toggles to auto-regenerate CLI tool profiles after a provider model sync. When enabled, a model-catalog change (re)writes that tool's profile files from the live catalog — Codex (`~/.codex/*.config.toml`) and now **Claude Code** (`~/.claude/profiles//settings.json`, via an extracted `syncClaudeProfilesFromModels` + a new `claudeProfileAutoSync.ts` mirroring the Codex path). Both are **off by default** and never touch the active/default CLI config; they are backed by the `OMNIROUTE_AUTO_SYNC_CODEX_PROFILES` / `OMNIROUTE_AUTO_SYNC_CLAUDE_PROFILES` feature flags (DB/dashboard override > env > default "false") and additionally gated behind the existing `CLI_ALLOW_CONFIG_WRITES` write-guard. A **"CLI profile auto-sync"** card on the CLI Code dashboard toggles each (moved from the providers dashboard in [#5778](https://github.com/diegosouzapw/OmniRoute/pull/5778) — thanks [@rdself](https://github.com/rdself)). Regression guards: `tests/unit/claude-profile-auto-sync-gate.test.ts`, `tests/unit/codex-profile-auto-sync-gate.test.ts`, `tests/unit/cli/setup-claude.test.ts` (follow-up to #5737). - **cli (startup banner):** the `serve` startup banner now prints the running OmniRoute version (`v3.8.x`) beneath the ASCII logo, so the active version is visible at a glance without a separate `--version` call. Regression guard: `tests/unit/cli-serve-version-banner.test.ts`. Thanks [@chirag127](https://github.com/chirag127) ([#5752](https://github.com/diegosouzapw/OmniRoute/pull/5752)). - **analytics (subscription cost):** flat-rate providers now show **$0** in cost analytics instead of an inflated per-token estimate. Subscription / coding-plan providers (every cookie-web provider — ChatGPT Web, grok-web, … — plus the dedicated **Minimax Coding**, **Kimi Coding**, **GLM Coding**, **Alibaba Coding Plan**, and **Xiaomi MiMo** plans) bill a flat fee, not per token, yet stil _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.43]_ ### v3.8.42 - Date: 2026-06-30 - Version: v3.8.42 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.42 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.42 - **added** — Add an inflation guard to the stacked compression pipeline that discards compressed output if token count is not reduced and sends the original request instead, with a pipeline-inflation-guard warning recorded in compression stats - **added** — Complete German, French, and Japanese compression rule packs with dedup and ultra categories for collapsing repeated context and abbreviating technical vocabulary - **added** — Add Chinese (zh / wenyan) input-side compression rule pack with repeated context collapsing, pleasantry dropping, and technical term abbreviation - **added** — Auto-detect Chinese compression language by distinguishing zh from ja through Han-without-kana analysis - **added** — Add Gradle and .NET CLI (dotnet) to the RTK tool-output filter catalog with recognition of build output and preservation of build status and error messages - **fixed** — Fix chatgpt-web provider failures on Electron desktop app by implementing a runtime-portable SHA-3 helper that falls back to pure-JS Keccak-f[1600] when native SHA-3 is unavailable in BoringSSL - **fixed** — Fix Bytez provider key validation by correcting the registry base URL to include the full OpenAI-compatibility chat path and replacing chat probe validation with a dedicated auth-only endpoint probe ###### ✨ New Features - **compression (pipeline):** add an honest default-on **inflation guard** to the stacked compression pipeline (T02 / Headroom H1). If the fully-stacked engines produce a body that did not actually shrink — its token count is `>=` the original — the compressed body is discarded and the verbatim original request is sent upstream instead, with a `pipeline-inflation-guard` warning recorded in the compression stats. This is safe by construction (the only fallback is the unmodified original, always a valid payload) and complements the existing opt-in per-step TV1 bail-out, which governs step-to-step advancement rather than the final output. New `open-sse/services/compression/pipelineGuards.ts`; wired at the single `finalizeStackedResult` choke point shared by the sync and async stacked paths. Regression guards (incl. an inflating-engine integration test) in `tests/unit/compression-pipeline-inflation-guard.test.ts`. - **compression (caveman):** complete the German, French, and Japanese rule packs with the `dedup` (repeated-context collapsing) and `ultra` (abbreviation / terse) categories they were missing — these three languages previously shipped only `context`/`filler`/`structural`, while `en`/`es`/`id`/`pt-BR` had all five. So a de/fr/ja conversation compressed at higher intensities now collapses repeated boilerplate ("wie bereits besprochen" → "Siehe oben.", "comme mentionné précédemment" → "Voir ci-dessus.", "前述のとおり" → "(上記参照)") and abbreviates dense technical vocabulary (`Datenbank`→`DB`, `Authentifizierung`→`Auth`; `base de données`→`BD`, `authentification`→`auth`; `データベース`→`DB`, `アプリケーション`→`app`). Patterns mirror the existing `es` pack and stay ReDoS-safe (bounded literal alternations; the CJK pack uses no `\b` since Japanese has no word boundaries). Regression guard: `tests/unit/caveman-packs-de-fr-ja.test.ts` (packs load + validate + shrink a representative sample). gaps v3.8.42 — T05/C2. - **compression (caveman):** add a **Chinese (zh / wenyan 文言) input-side rule pack** — the counterpart of the existing output-side `terse-cjk` style. New `rules/zh/{dedup,filler,ultra}.json` collapse repeated context ("如前所述" → "见上。"), drop pleasantries/hedging ("请帮我…/谢谢/我觉得"), strip sentence-final modal particles ("吗/呢/吧"), and abbreviate dense technical terms ("数据库"→"DB", "应用程序"→"app"). Chinese is now auto-detected: `detectCompressionLanguage` distinguishes zh from ja by Han-without-kana (kana is Japanese-exclusive, so a Han-heavy Japanese sentence still resolves to `ja`), and `zh` is listed in `listSupportedCompressionLanguages`. Patterns are ReDoS-safe (bounded literal alternations, no `\b` since CJK has no word boundaries). Regression guard: `tests/unit/caveman-packs-zh-wenyan.test.ts` (packs load + validate + shrink; zh/ja/non-CJK detection). gaps v3.8.42 — T05/C6. - **compression (RTK):** add **Gradle** and **.NET CLI (`dotnet`)** to the RTK tool-output filter catalog. Tool output for `gradle`/`gradlew` and `dotnet build|test|restore|publish` is now recognized (both by command and by output content) and compressed: Gradle daemon/welcome banners and no-op `> Task … UP-TO-DATE/SKIPPED/FROM-CACHE` lines are dropped while `BUILD SUCCESSFUL/FAILED`, "What went wrong", and stack traces are preserved; the .NET build banner, copyright, and `Determining projects to restore`/`Restored …` chatter are dropped while `Build succeeded/FAILED`, `error CS####`/`warning CS####`, and test summaries are preserved. New builtin filters `engines/rtk/filters/{gradle,dotnet}.json` (with inline tests run by the catalog gate) plus `gradle`/`dotnet` entries in the command detector. Regression guard: `tests/unit/rtk-gradle-dotnet-filters.test.ts`. gaps v3.8.42 — T07/R9. ###### 🔧 Bug Fixes - **providers (chatgpt-web):** fix `502 ChatGPT sentinel failed: Digest method not supported` on the **Electron desktop app**, which made every `chatgpt-web/*` request fail. The sentinel proof-of-work hashed with native `createHash("sha3-512")`, bu _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.42]_ ### v3.8.41 - Date: 2026-06-29 - Version: v3.8.41 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.41 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.41 - **added** — Add selectable relay backend (TS / Bifrost / auto) via OMNIROUTE_RELAY_BACKEND / RELAY_ROUTING_BACKEND environment variable, with auto mode selecting Bifrost when BIFROST_BASE_URL is set and falling back to TypeScript if the sidecar is unreachable - **added** — Add X-Routing-Backend and X-Routing-Fallback response headers to indicate which relay backend handled the request - **added** — Add Proxy Pool dropdown to OpenCode Free per-account proxy modal, allowing selection of pre-saved proxies instead of manual Host/Port/credentials entry - **added** — Add Saved / Custom toggle for proxy configuration in OpenCode, with server-side resolution of proxies from the registry - **fixed** — Fix Claude translator to synthesize a minimal user turn when an OpenAI request carries only system/developer messages, preventing 400 errors from the Messages API - **fixed** — Remove retired Google AI Studio model IDs and align the Gemini catalog to live GenAI API, removing gemini-1.5-pro, gemini-1.5-flash, gemini-2.0-flash, gemini-2.0-flash-lite, and adding gemini-3.1-flash-lite and gemini-embedding-001/gemini-embedding-2 - **fixed** — Add modelDeprecation forwards for renamed Gemini model IDs to redirect legacy IDs to GA models instead of returning 404 - **fixed** — Fix embedded-services dashboard by lazily initializing service supervisors from /api/services/[name]/logs to prevent 404 errors for cliproxy and 9router logs before bootstrap - **fixed** — Fix dashboard lifecycle buttons to send JSON and properly handle empty install bodies defaulting to version latest - **fixed** — Fix dashboard lifecycle and log-stream failures to surface as actionable UI errors instead of silently showing no logs - **fixed** — Fix dashboard to recognize Tailscale CGNAT 100.64.0.0/10 peers as private-LAN local for local-only service access - **fixed** — Fix /api/v1/providers/{cliproxyapi,9router}/models to return synced embedded-service models instead of invalid_provider - **fixed** — Fix Claude adaptive-thinking dashboard Thinking-Budget setting being dropped on restart by hydrating from settings in server-init - **fixed** — Fix Claude executor to honor mode auto for stripping adaptive thinking while keeping default passthrough behavior, and remap operator thinking.type enabled to adaptive shape - **fixed** — Fix Claude thinking block replay to emit signature-less redacted_thinking block instead of reconstructing reasoning_content with fabricated signature - **fixed** — Fix OpenCode proxy resolution to safely degrade when proxyId is unknown or deleted by falling back to direct connection ##### [3.8.41] — 2026-06-29 ###### ✨ New Features - **feat(relay): selectable relay backend (TS / Bifrost / `auto`)** — the OpenAI-compatible relay endpoint can now route its hot path through a native Bifrost sidecar without clients changing URLs. `OMNIROUTE_RELAY_BACKEND` / `RELAY_ROUTING_BACKEND` = `ts | bifrost | auto`: defaults to the existing TypeScript relay; `auto` selects Bifrost when `BIFROST_BASE_URL` is set (and `BIFROST_ENABLED` ≠ `0`) and falls back to TS automatically if the sidecar is unreachable; `bifrost` keeps strict failure behavior. Auth, per-IP/token rate limits, prompt-injection checks, and model allowlists still run in the Next relay route before dispatch (control plane stays in the app); responses carry `X-Routing-Backend` / `X-Routing-Fallback`. Regression guards: `tests/unit/api/v1/relay-routing-backend.test.ts`, `tests/unit/api/v1/bifrost-sidecar.test.ts`. ([#5315](https://github.com/diegosouzapw/OmniRoute/pull/5315), #5316 — thanks @KooshaPari) ###### 🔧 Bug Fixes - **translator (claude):** synthesize a minimal `user` turn when an OpenAI→Claude request carries **only** `system`/`developer` messages, so the request stops failing with `[400]: messages: at least one message is required`. `openaiToClaudeRequest` hoists every system/developer turn into Claude's top-level `system` field and filters them out of `messages`; an all-system input (OpenCode compaction / title-generation requests) left `messages: []`, which the Messages API rejects — surfacing in OpenCode as a mid-task `stream error` that drops the conversation. The guard fires only when `messages` would otherwise be empty (system instructions still drive the response), so non-empty requests are unaffected. ([#5342](https://github.com/diegosouzapw/OmniRoute/pull/5342) — thanks @wild-feather) - **providers (gemini):** drop retired Google AI Studio model ids and align the catalog to what the live GenAI API actually serves (verified 2026-06-29 against the official deprecations page). Removes long-retired `gemini-1.5-pro`/`gemini-1.5-flash`, the shut-down `gemini-2.0-flash`/`gemini-2.0-flash-lite`, and dead experimentals; renames `gemini-3.1-flash-lite-preview` → the GA `gemini-3.1-flash-lite`; swaps the retired `text-embedding-004` for the live `gemini-embedding-001`/`gemini-embedding-2`; and adds graceful `modelDeprecation` forwards so legacy/renamed ids redirect to the GA model instead of 404ing. Native AI-Studio-direct image/video/music registration is intentionally out of scope (needs real executor work; those models stay reachable via Antigravity/Vertex/aggregators). ([#5337](https://github.com/diegosouzapw/OmniRoute/pull/5337) — thanks @backryun) - **services (dashboard):** fix the embedded-services dashboard failures (#5298) — service supervisors are now lazily initialized from `/api/services/[name]/logs` so `cliproxy`/`9router` logs no longer 404 before bootstrap registers a supervisor; lifecycle buttons send JSON (empty install bodies default to `version: "latest"`, malformed JSON still returns `400 Invalid JSON body`); lifecycle and log-stream failures surface as actionable UI errors instead of silently showing no logs; Tailscale CGNAT `100.64.0.0/10` peers count as private-LAN local for local-only service access; a parent `/dashboard/context` → `/dashboard/context/settings` redirect stops RSC prefetch 404s; and `/api/v1/providers/{cliproxyapi,9router}/models` return synced embedded-service models instead of `invalid_provider`. ([#5299](https://github.com/diegosouzapw/OmniRoute/pull/5299), #5298 — thanks @KooshaPari) - **thinking (claude):** fix three independent defects in Claude adaptive-thinking on the OpenAI-compatible path (Cursor → Claude OAuth). **(A)** the dashboard Thinking-Budget setting was dropped on every restart — `setThinkingBudgetConfig` was never called at boot, so a saved `{mode:"adaptive"…}` silently reverted to passthrough; it's now hydrated from settings in `server-init`. **(B)** the Claude executor force-injected _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.41]_ ### v3.8.40 - Date: 2026-06-29 - Version: v3.8.40 - Original notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.40 - Permalink: https://whatsnew.fyi/product/omniroute/releases/v3.8.40 - **added** — Add relevance extractive compression engine that scores sentences by term-overlap with user queries and greedily keeps the most relevant within a budget - **added** — Add hard-budget compression mode that trims bodies to a token budget by ranking and dropping lowest-saliency sentences while preserving original order - **added** — Add opt-in result memoization for deterministic compression engines to skip recompute on the hot path - **added** — Extend X-OmniRoute-Compression response header to surface compression transparency annotation with token counts and rules applied - **added** — Add saliency heatmap visualization in the compression studio to color tokens by saliency score - **added** — Add composite-command splitter for RTK detection to recognize individual commands in chains like cd /x && git status - **added** — Add omniroute_tool_search MCP tool for lexical keyword search over MCP tool names and descriptions - **added** — Add opt-in RTK semantic command-output renderers that rewrite structured command output into compact forms - **changed** — Extend X-OmniRoute-Compression header to append compression statistics while maintaining backward compatibility with existing parsers ##### [3.8.40] — TBD _In development — bullets added per PR; finalized at release._ ###### ✨ New Features - **feat(compression): relevance extractive engine** — a new opt-in compression engine that scores each sentence by term-overlap (Jaccard) with the user's last query minus a length/boilerplate penalty, greedily keeps the most relevant within a budget, and reconstructs the original order. Pure-string, deterministic, ReDoS-safe (char-code tokenization, no `RegExp` over user input), fail-open, default off. Ideal for trimming long pasted RAG context / tool output to what's relevant. Sentences carrying real signal (digits/URLs/errors/code/paths) are never dropped; `overlapThreshold`/`budgetPercent`/`boilerplateWeight` are configurable. Tier-2 item of the compression feature-extraction roadmap (#7). ([#5289](https://github.com/diegosouzapw/OmniRoute/pull/5289)) - **feat(compression): hard-budget mode — compress to ≤ N tokens** — a deterministic post-pass (`targetTokens` / `targetRatio`, default unset → no-op) that trims a body to a token budget. It ranks sentences/lines by average `scoreToken` ascending and drops the lowest-saliency ones until the body fits (measured by the exact cl100k `countTextTokens`), preserving original order. Lines carrying real signal (digits, URLs, `Error:`-family, code fences, stack `at`-frames, multi-segment paths, `key=value`) are never dropped; the budget is distributed proportionally across messages so the total stays ≤ target; an unreachable target (all-preserved) surfaces a `validationWarnings` note instead of failing silently. Does NOT touch the `estimateCompressionTokens` budget-gate estimator. Tier-3 item of the compression feature-extraction roadmap (#17). ([#5288](https://github.com/diegosouzapw/OmniRoute/pull/5288), follow-up [#5291](https://github.com/diegosouzapw/OmniRoute/pull/5291)) - **feat(compression): result memoization for deterministic engines (opt-in)** — caches `(input, config) → result` for provably pure, stateless modes (`lite`/`standard`/`rtk` and stacked pipelines of `{lite,caveman,rtk}`) to skip recompute on the hot path. Opt-in via `memoizeCompressionResults` (default off → zero behavior change). Conservative opt-in whitelist (stateful `ccr`/`session-dedup` — which write the cross-request CCR store — and model-backed `ultra`/`aggressive`/`llmlingua` are never cached), principal-scoped (skipped without a principal, so no cross-principal body leak), and clone-on-store + clone-on-read. Tier-3 item of the compression feature-extraction roadmap (#21). ([#5286](https://github.com/diegosouzapw/OmniRoute/pull/5286)) - **feat(compression): inline transparency annotation** — surfaces `tokens=847→312; rules: filler×8, dedup×2` derived from existing compression stats. The `X-OmniRoute-Compression` response header is extended **append-only** (the `mode; source=X` prefix stays byte-identical, so existing header parsers don't break) and the compression studio cockpit shows a matching badge. Zero new computation — it aggregates the `rulesApplied`/`techniquesUsed` already on the stats. Tier-3 item of the compression feature-extraction roadmap (#18). ([#5284](https://github.com/diegosouzapw/OmniRoute/pull/5284)) - **feat(compression): saliency heatmap in the compression studio** — the preview studio can now color each token by saliency: `ultra` per-token `scoreToken` (0–1, green→red gradient) or universal kept/removed from the existing diff. A dry-run visualization behind a toggle (no cost on a normal preview; backward-compatible when off). Completes the visualization half of roadmap item #13 (the A/B comparison shipped in [#5080](https://github.com/diegosouzapw/OmniRoute/pull/5080)). ([#5285](https://github.com/diegosouzapw/OmniRoute/pull/5285)) - **feat(compression): composite-command splitter for RTK detection** — `cd /x && git status` now detects as `git-status` (previously the whole string was treated as one command and matched no filter). A quote-aware top-level tokenizer splits on _[Truncated at 4000 characters — full notes: https://github.com/diegosouzapw/OmniRoute/releases/tag/v3.8.40]_