Compare commits

..

87 Commits

Author SHA1 Message Date
diegosouzapw e47740e02e feat: sub2api T05/T08/T09/T13/T14 + bump to 3.0.0-rc.7
Build Electron Desktop App / Validate version (push) Failing after 34s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
2026-03-22 23:17:52 -03:00
diegosouzapw d9ff0035f5 chore: bump version to 3.0.0-rc.6 (sub2api gap tasks T01-T15)
Build Electron Desktop App / Validate version (push) Failing after 29s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
2026-03-22 21:01:33 -03:00
diegosouzapw 7a7f3be0d2 feat(sub2api): implement T01-T15 gap analysis tasks (3.0.0-rc.6)
T01 (P1): requested_model column in call_logs
- Migration 009_requested_model.sql: ALTER TABLE call_logs ADD COLUMN requested_model
- callLogs.ts: INSERT + SELECT updated to include requestedModel field

T02 (P1): Strip empty text blocks from nested tool_result.content
- New stripEmptyTextBlocks() recursive helper in openai-to-claude.ts
- Applied on tool_result content before forwarding to Anthropic
- Prevents 400 'text content blocks must be non-empty' errors

T03 (P1): Parse x-codex-5h-*/x-codex-7d-* headers for precise quota reset
- parseCodexQuotaHeaders() in codex.ts extracts usage/limit/resetAt
- getCodexResetTime() returns furthest-out reset timestamp for safe unblocking

T04 (P1): X-Session-Id header for external sticky routing
- extractExternalSessionId() in sessionManager.ts reads x-session-id,
  x-omniroute-session, session-id headers with 'ext:' prefix to avoid collisions

T06 (P2): account_deactivated permanent expired status on 401
- ACCOUNT_DEACTIVATED_SIGNALS constant + isAccountDeactivated() in accountFallback.ts
- Returns 1-year cooldown (effectively permanent) to prevent retrying dead accounts

T07 (P2): X-Forwarded-For IP validation
- New src/lib/ipUtils.ts with extractClientIp() and getClientIpFromRequest()
- Skips 'unknown'/non-IP entries in X-Forwarded-For chain

T10 (P2): credits_exhausted distinct account status
- CREDITS_EXHAUSTED_SIGNALS + isCreditsExhausted() in accountFallback.ts
- Returns 1h cooldown with creditsExhausted flag, distinct from rate_limit 429

T11 (P1): max reasoning_effort -> budget_tokens: 131072
- EFFORT_BUDGETS and THINKING_LEVEL_MAP updated with max: 131072, xhigh: 131072
- Reverse mapping now returns 'max' for full-budget responses
- Unit test updated to expect 'max' (was 'high')

T12 (P3): Model pricing updates
- MiniMax M2.7 / MiniMax-M2.7 / minimax-m2.7-highspeed pricing added

T15 (P1): Array content normalization for system/tool messages
- normalizeContentToString() helper exported from openai-to-claude.ts
- System messages with array content now correctly collapsed to string
2026-03-22 20:55:35 -03:00
diegosouzapw 91e45fbe95 chore: remover new-features-sub21 do tracking do git
Remover as exceções !docs/new-features-sub21/ do .gitignore para que
a pasta de tasks internas não seja mais rastreada pelo git.
2026-03-22 20:32:17 -03:00
diegosouzapw 7d7e9da28c feat(providers): adicionar provedor Puter AI com 500+ modelos
Registrar o provedor Puter como gateway OpenAI-compatible que expõe
modelos de múltiplos fornecedores (GPT, Claude, Gemini, Grok, DeepSeek,
Qwen, Mistral, Llama) através de um único endpoint REST.

- Criar PuterExecutor com autenticação Bearer token
- Adicionar entrada no providerRegistry com 40+ modelos curados
- Habilitar passthroughModels para acesso aos 500+ modelos do catálogo
- Registrar alias "pu" para acesso rápido
- Adicionar metadados do provedor em shared/constants/providers.ts
2026-03-22 20:29:06 -03:00
diegosouzapw 24a9739604 docs: add sub2api gap analysis + 15 implementation tasks
Add competitive analysis of sub2api (v0.1.104, 87 contributors)
comparing features, open PRs, and model pricing against OmniRoute.

Files:
- docs/new-features-sub21/gap-analysis.md — full analysis (commits + 38 open PRs)
- docs/new-features-sub21/implementation-plan.md — phased plan for all 15 gaps
- docs/new-features-sub21/tasks/T01-T15 — detailed task files with:
  - Problem description + sub2api PR references
  - Step-by-step implementation with code snippets
  - Affected files list
  - Acceptance criteria

Priority breakdown:
  P1 (4): requested_model logs, empty tool_result blocks, x-codex-* headers, X-Session-Id
  P2 (6): rate-limit persistence, account_deactivated, XFF validation, session limits, Codex/Spark scopes, credits_exhausted
  P3 (5): max reasoning effort, model pricing, stale quota display, proxy fast-fail, array content

Source: https://github.com/Wei-Shaw/sub2api
2026-03-22 18:12:50 -03:00
diegosouzapw 4fb9687782 docs(3.0.0-rc.5): comprehensive CHANGELOG and README vs v2.9.5
Build Electron Desktop App / Validate version (push) Failing after 31s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
- CHANGELOG: [3.0.0-rc.5] section now serves as full 'What's New vs v2.9.5':
  * 2 new providers (OpenCode Zen/Go via PR #530)
  * 3 new features: Registered Keys API (#464), provider icons (#529), model auto-sync (#488)
  * 10 bug fixes (#521, #522, #524, #527, #532, #535, #536, #537, #489, #510, #492)
  * 16 issues resolved total, DB migration 008
- README: added 'What's New in v3.0.0' table section after badges
2026-03-22 15:51:54 -03:00
diegosouzapw 95ffc21b60 feat(3.0.0-rc.5): Registered Keys Provisioning API (#464)
Complete implementation of auto-provisioning API:
- DB migration 008: registered_keys, provider_key_limits, account_key_limits
- src/lib/db/registeredKeys.ts: full quota enforcement, idempotency, sha256
  hashing, budget tracking, window auto-reset
- POST /api/v1/registered-keys — issue with quota check
- GET /api/v1/registered-keys — list (masked)
- GET|DELETE /api/v1/registered-keys/[id] — get/revoke
- POST /api/v1/registered-keys/[id]/revoke — explicit revoke
- GET /api/v1/quotas/check — pre-validate without issuing
- GET|PUT /api/v1/providers/[id]/limits — provider limits CRUD
- GET|PUT /api/v1/accounts/[id]/limits — account limits CRUD
- POST /api/v1/issues/report — optional GitHub issue reporting
  (requires GITHUB_ISSUES_REPO + GITHUB_ISSUES_TOKEN env vars)
- Exported all from localDb.ts
2026-03-22 15:33:45 -03:00
diegosouzapw f3c5e55b26 feat(3.0.0-rc.4): merge PR #530 — OpenCode Zen and Go providers
Build Electron Desktop App / Validate version (push) Failing after 39s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
Includes all commits from @kang-heewon's PR #530:
- OpencodeExecutor with multi-format routing
- opencode-zen + opencode-go registered in provider registry
- UI metadata added to providers.ts
- Unit tests for OpencodeExecutor (improved to avoid state coupling)

Cherry-picked from add-opencode-providers into 3.0.0-rc.
Conflicts resolved: executors/index.ts (merged pollinations+cloudflare-ai),
providerRegistry.ts (kept testKeyBaseUrl from rc.2 + PR's authType/models).
2026-03-22 15:23:00 -03:00
kang-heewon 40183c6a5c test(providers): improve OpencodeExecutor tests to avoid internal state coupling 2026-03-22 15:22:38 -03:00
kang-heewon 457c59e38a test(providers): add unit tests for OpencodeExecutor 2026-03-22 15:22:38 -03:00
diegosouzapw aa93a3f2e2 feat(3.0.0-rc.3): provider icons, model auto-sync, Gemini OAuth fix
Build Electron Desktop App / Validate version (push) Failing after 40s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
feat(ui): ProviderIcon component with @lobehub/icons + PNG fallback (#529)
  - 130+ providers covered by Lobehub SVG components via LobehubErrorBoundary
  - Falls back to existing /providers/{id}.png, then generic icon
  - Replaces manual img state machine in ProviderCard + ApiKeyProviderCard

feat(scheduler): modelSyncScheduler — 24h model list auto-update (#488)
  - Syncs 16 major providers every 24h (MODEL_SYNC_INTERVAL_HOURS configurable)
  - Wired into POST /api/sync/initialize startup hook

fix(oauth): Gemini CLI — clear error when client_secret missing in Docker (#537)
2026-03-22 15:01:38 -03:00
diegosouzapw 8b9abcb6cc fix(3.0.0-rc.2): resolve issues #536, #535, #524
Build Electron Desktop App / Validate version (push) Failing after 34s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
fix(providers): LongCat AI key validation — correct base URL and auth header (#536)
  - baseUrl: longcat.chat/api/v1/chat/completions -> api.longcat.chat/openai
  - authHeader: 'bearer' -> 'Authorization' + authPrefix: 'Bearer'

fix(combo): implement pinnedModel override in comboAgentMiddleware (#535)
  - Previously: pinnedModel was detected but body.model was never updated
  - Now: body = { ...body, model: pinnedModel } when context_cache_protection fires

fix(cli-tools): add OpenCode config save to guide-settings endpoint (#524)
  - Added 'opencode' case to switch in guide-settings/[toolId]/route.ts
  - saveOpenCodeConfig(): XDG_CONFIG_HOME aware, writes [provider.omniroute] TOML block
2026-03-22 13:31:56 -03:00
diegosouzapw 1ecc1908c7 chore(3.0.0-rc.1): bump version to 3.0.0-rc.1, close resolved issues, update CHANGELOG
Build Electron Desktop App / Validate version (push) Failing after 30s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
- package.json: 2.9.5 → 3.0.0-rc.1
- docs/openapi.yaml: version → 3.0.0-rc.1
- CHANGELOG.md: add [3.0.0-rc.1] section with all batch1-3 fixes
- scripts/check-docs-sync.mjs: isSemver now accepts pre-release versions (X.Y.Z-prerelease.N)

Closed issues: #489, #492, #510, #513, #520, #521, #522, #525, #527, #532
RC versioning: rc.1 → rc.2 → rc.N on each VPS deploy until v3.0.0 is approved
2026-03-22 12:25:30 -03:00
diegosouzapw 6a2c7b467d fix(3.0.0-rc/batch3): convert tool_result blocks to text to stop Codex loop (#527)
fix(chat): convert tool_result content blocks to [Tool Result: id] text (#527)
  - Previously, tool_result blocks in user messages were silently dropped
  - This caused an infinite loop when Claude Code + superpowers routed to Codex:
    Codex never received the tool response and kept re-requesting the tool
  - Now: tool_result → text block '[Tool Result: {id}]\n{content}'
  - Handles string, array-of-text, and JSON-serialized content types

docs(issues): add Turbopack postinstall workaround on #509 and #508
docs(issues): note that #464 (API key provisioning) is on the v3.0 roadmap
2026-03-22 11:47:39 -03:00
diegosouzapw 0acef57865 fix(3.0.0-rc/batch2): resolve issues #510, #492, and improve #520, #529
fix(cli): normalize MSYS2/Git-Bash paths in cliRuntime.ts (#510)
  - Add normalizeMsys2Path() helper: /c/Program Files/... → C:\Program Files\...
  - Apply to both Windows 'where' and Unix 'command -v' path resolution
  - Fixes 'CLI not detected' on Windows when running Git Bash / MSYS2

fix(cli-launcher): detect mise/nvm on server.js not found error (#492)
  - Show targeted fix instructions based on which Node manager is in use
  - mise users: told to use npx or mise exec
  - nvm users: reminded to nvm use --lts before reinstalling

docs(issues): add pnpm bindings workaround comment (#520)
docs(issues): note OpenCode/Lobehub icons coming in v3.0.0 (#529)
2026-03-22 11:41:04 -03:00
diegosouzapw 43046ee649 fix(3.0.0-rc/batch1): resolve issues #521, #522, #525, #532, #489
fix(login): redirect to /dashboard/onboarding when API returns needsSetup:true (#521)
  - Handle the case where user skips password setup and lands on login
  - Instead of showing a cryptic error, redirect to onboarding flow

fix(api-manager): replace useless 'copy masked key' button with lock tooltip (#522)
  - Copying a masked key (sk-proj123****abcd) is misleading and useless
  - Show a lock icon on hover explaining key is only available at creation time
  - Add i18n key 'keyOnlyAvailableAtCreation'

fix(opencode-go): use zen/v1 for API key validation, not zen/go/v1 (#532)
  - Added testKeyBaseUrl field to RegistryEntry interface
  - opencode-go: testKeyBaseUrl → zen/v1 (same key authenticates both tiers)
  - validation.ts: resolveBaseUrl for key testing now prefers testKeyBaseUrl

fix(antigravity): return structured 422 error when projectId is missing (#489)
  - Instead of throwing (crash), executor returns an OpenAI-format error JSON
  - Client receives message with instruction to reconnect OAuth
  - Prevents opaque 500 errors in the proxy logs

chore: close #525 (OmniRoute = 9router — same project, different name)
docs: add Docker password reset comment on #513 with INITIAL_PASSWORD workaround
2026-03-22 11:31:34 -03:00
Diego Rodrigues de Sa e Souza a15fda0c08 Merge pull request #534 from diegosouzapw/release/v2.9.5
Build Electron Desktop App / Validate version (push) Failing after 33s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
chore(release): v2.9.5 — OpenCode providers, embedding fix, CLI masked key fix
2026-03-22 10:32:33 -03:00
diegosouzapw e5988764ce chore(release): v2.9.5 — OpenCode providers, embedding credentials fix, CLI masked key fix, CACHE_TAG_PATTERN fix
- feat(providers): add OpenCode Zen and Go providers with multi-format executor (PR #530 by @kang-heewon)
- fix(embeddings): use provider node ID for custom embedding provider credential lookup (PR #528 by @jacob2826)
- fix(cli-tools): resolve real API key from DB (keyId) before writing to CLI config files (#523, #526)
- fix(combo): update CACHE_TAG_PATTERN to match literal \\n prefix/suffix around omniModel tag (#531)
- chore: bump version to 2.9.5 in package.json + docs/openapi.yaml
- docs: update CHANGELOG.md with v2.9.5 release notes
2026-03-22 10:30:04 -03:00
diegosouzapw 9c9d9b5a8d feat(providers): add OpenCode Zen and Go providers (#530) 2026-03-22 10:25:15 -03:00
kang-heewon 44dc564d85 chore: remove GHCR workflow from upstream PR 2026-03-22 10:24:50 -03:00
kang-heewon 83e367afab ci: add GHCR publish workflow for fork deployments 2026-03-22 10:24:50 -03:00
kang-heewon 8b7e7c2669 test(providers): improve OpencodeExecutor tests to avoid internal state coupling 2026-03-22 10:24:50 -03:00
kang-heewon 53474021b7 test(providers): add unit tests for OpencodeExecutor 2026-03-22 10:24:50 -03:00
kang-heewon da1ed1b5b2 feat(providers): register opencode-zen and opencode-go in provider registry 2026-03-22 10:24:50 -03:00
kang-heewon e08d661600 feat(providers): register opencode executors and add UI metadata
- Register OpencodeExecutor for 'opencode-zen' and 'opencode-go' in executors map
- Add OpencodeExecutor export in index.ts
- Add UI metadata for both providers in APIKEY_PROVIDERS:
  - OpenCode Zen: https://opencode.ai/zen
  - OpenCode Go: https://opencode.ai/zen/go
- Both use 'opencode' icon with #6366f1 color
2026-03-22 10:24:50 -03:00
kang-heewon 1aa1bc7a26 feat(providers): add OpencodeExecutor for opencode-zen/go multi-format routing 2026-03-22 10:23:32 -03:00
Diego Rodrigues de Sa e Souza 47634e942e Merge pull request #533 from diegosouzapw/fix/issues-521-523-526-531
fix: resolve masked key in CLI config saves + CACHE_TAG_PATTERN \n handling (#523, #526, #531)
2026-03-22 10:23:19 -03:00
Diego Rodrigues de Sa e Souza 15466cbf1a Merge pull request #528 from jacob2826/codex/fix-embedding-compatible-provider-credentials
fix: use provider node credentials for custom embedding providers
2026-03-22 10:23:16 -03:00
diegosouzapw 2a749db427 fix: resolve masked key bug in CLI config saves, fix CACHE_TAG_PATTERN for \n prefix (#523, #526, #531)
fix(cli-tools): save real API key to CLI config files instead of masked string (#523, #526)
  - claude-settings/route.ts: accept keyId, look up real key from DB (getApiKeyById)
  - cline-settings/route.ts: same keyId resolution pattern
  - openclaw-settings/route.ts: same keyId resolution pattern
  - ClaudeToolCard.tsx: store key.id as selected value, send keyId in POST body
  The /api/keys endpoint returns masked strings (first8+****+last4) which were being
  written verbatim to ~/.claude/settings.json and similar config files, causing auth
  failures on CLI tool launch.

fix(combo): update CACHE_TAG_PATTERN to strip surrounding \\n sequences (#531)
  - comboAgentMiddleware.ts: non-global regex now matches literal \\n (backslash-n)
    and actual newline U+000A that combo.ts injects around the <omniModel> tag.
2026-03-22 09:49:03 -03:00
jacob2826 ecccce86e4 fix: use provider node credentials for embeddings 2026-03-22 16:22:58 +08:00
Diego Rodrigues de Sa e Souza bf3f64bea4 Merge pull request #519 from diegosouzapw/release/2.9.4
Build Electron Desktop App / Validate version (push) Failing after 32s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
chore(release): v2.9.4 — bug fixes (#491, #515, #517)
2026-03-21 17:40:23 -03:00
diegosouzapw 2f2d6b8535 chore(release): v2.9.4 — bug fixes (#491, #515, #517)
- fix(translator): preserve prompt_cache_key in Responses API translation (#517)
- fix(combo): escape \n in tagContent for valid JSON injection (#515)
- fix(usage): sync expired token status back to DB on live auth failure (#491)
- chore: bump version to 2.9.4 in package.json + docs/openapi.yaml
- docs: update CHANGELOG.md with v2.9.4 release notes
2026-03-21 17:37:51 -03:00
Diego Rodrigues de Sa e Souza d68c884649 Merge pull request #518 from diegosouzapw/fix/issue-517-515-prompt-cache-key-tagcontent
fix: preserve prompt_cache_key in Responses API, escape \n in tagContent (#517, #515)
2026-03-21 17:32:24 -03:00
diegosouzapw 8b556de03b fix: preserve prompt_cache_key in Responses API translation, escape \n in tagContent (#517, #515)
fix(translator): preserve prompt_cache_key when translating Responses API requests
  (#517) — prompt_cache_key is an account-affinity signal used by Codex for
  prompt cache routing. Deleting it from the translated request prevented full
  cache effectiveness. Removed delete from openai-responses.ts and
  responsesApiHelper.ts cleanup blocks.

fix(combo): escape \n in tagContent so injected JSON string is valid (#515)
  — omniModel tag content used template literal newlines (U+000A) which produce
  unescaped newline chars inside a JSON string value. Replaced with literal \n
  escape sequences for valid JSON injection in streaming SSE content chunks.
2026-03-21 17:09:13 -03:00
Diego Rodrigues de Sa e Souza 7229af53c3 Merge pull request #516 from diegosouzapw/release/2.9.3
Build Electron Desktop App / Validate version (push) Failing after 25s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
feat(providers): 5 new free AI providers — v2.9.3
2026-03-21 16:55:29 -03:00
diegosouzapw 81b3034c2f feat(providers/logos): add logos for 5 new free providers
- public/providers/longcat.png — pink cat icon (generated)
- public/providers/pollinations.png — pixel bee icon (generated)
- public/providers/aimlapi.png — indigo neural network icon (generated)
- public/providers/cloudflare-ai.svg — Cloudflare official SVG (simpleicons.org)
- public/providers/scaleway.svg — Scaleway official SVG (simpleicons.org)

Icons serve at /providers/{id}.png (PNG fallback to SVG)
2026-03-21 16:47:49 -03:00
diegosouzapw f0419396b5 chore(release): bump version to 2.9.3, update CHANGELOG
- Version bumped from 2.9.2 → 2.9.3 in package.json + docs/openapi.yaml
- CHANGELOG.md updated with full release notes for 2.9.3
  (5 new free providers, 2 metadata updates, 2 custom executors, docs)
2026-03-21 15:44:35 -03:00
diegosouzapw 6b9c2754e8 feat(providers): add LongCat AI, Pollinations, Cloudflare AI, Scaleway, AI/ML API
New free providers:
- LongCat AI (lc/): 50M tokens/day free during public beta
- Pollinations AI (pol/): no API key needed, GPT-5/Claude/DeepSeek/Llama free
- Cloudflare Workers AI (cf/): 10K Neurons/day, ~150 LLM responses, Whisper free
- Scaleway AI (scw/): 1M free tokens for new accounts (EU/GDPR, Paris)
- AI/ML API (aiml/): $0.025/day credits, 200+ models via single endpoint

Provider metadata updates:
- Together AI: hasFree=true + 3 permanently free model IDs (Llama 70B, Vision, DeepSeek)
- Gemini: hasFree=true + freeNote (1,500 req/day free, no credit card)
- NVIDIA NIM: already had hasFree=true, confirmed correct

New executors:
- open-sse/executors/pollinations.ts: optional auth (no key support)
- open-sse/executors/cloudflare-ai.ts: dynamic URL with accountId credential

Documentation:
- README.md: 11-provider Ultimate Free Stack, 4 new pricing table rows
- README.md: LongCat/Pollinations/Cloudflare AI/Scaleway provider detail sections
- docs/i18n/pt-BR/README.md: updated pricing table + 4 new free provider sections
- docs/i18n/cs/README.md: combo stack updated

Tests: 821/821 pass (no regressions)
2026-03-21 15:40:05 -03:00
diegosouzapw 8edb131f8b docs: add npm downloads and Docker Hub pulls badges to README 2026-03-21 14:48:48 -03:00
Diego Rodrigues de Sa e Souza d6f6520a79 Merge pull request #514 from diegosouzapw/release/v2.9.2
Build Electron Desktop App / Validate version (push) Failing after 35s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
chore(release): v2.9.2 — Transcription Content-Type fix, Deepgram language detection, TTS error display
2026-03-21 14:03:33 -03:00
diegosouzapw cc2bb4d719 chore: update generate-release workflow to two-phase PR-first flow
Phase 1: bump, docs, i18n, commit, push, open PR → STOP for user confirmation
Phase 2 (post-merge): tag, GitHub release, Docker Hub, deploy both VPS
2026-03-21 13:58:08 -03:00
diegosouzapw 3859f1c9ae chore(release): v2.9.2 — transcription Content-Type fix, Deepgram language detection, TTS error display
- fix(transcription): resolveAudioContentType() maps video/mp4 → audio/mp4 for Deepgram/HuggingFace
- fix(transcription): detect_language=true + punctuate=true for Deepgram auto-detection
- fix(tts): upstreamErrorResponse() correctly extracts string from nested error objects
- docs: README transcription/TTS rows updated with provider counts and capabilities
- i18n: sync 29/30 language README files with updated feature descriptions
- chore: bump version 2.9.1 → 2.9.2
2026-03-21 13:54:22 -03:00
diegosouzapw 5f8d774e19 fix: [object Object] error display in TTS/transcription upstream errors
upstreamErrorResponse() now guards against parsed.error being an
object (e.g. ElevenLabs { error: { message, status_code } }) instead
of blindly using it as the error message string.
Both audioSpeech.ts and audioTranscription.ts fixed.
2026-03-21 10:47:55 -03:00
diegosouzapw 538a3e855c fix: transcription Content-Type + language detection for Deepgram/HuggingFace
- Add resolveAudioContentType() to map video/* MIME to audio/* (fixes .mp4 uploads returning 'no speech detected')
- Add detect_language=true for Deepgram auto-language detection (fixes non-English audio)
- Add punctuate=true for better output quality
- Forward language form param to Deepgram when provided
- Apply same Content-Type fix to HuggingFace handler
2026-03-21 10:38:57 -03:00
diegosouzapw 03f2ef1e2b fix: omniModel SSE tag data loss + v2.9.1 release (#511)
Build Electron Desktop App / Validate version (push) Failing after 38s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
2026-03-21 08:55:28 -03:00
Diego Rodrigues de Sa e Souza 237d0746cf Merge pull request #512 from zhangqiang8vip/feat/zws-v6
feat: per-protocol model compatibility, HMR leak fixes, and dev performance (V2-V5)
2026-03-21 08:53:54 -03:00
zhang-qiang 33b6c58087 fix(compat): store explicit false for per-protocol normalizeToolCallId
The truthy check treated false as falsy and deleted the property, preventing users from explicitly disabling normalization for a specific protocol when the top-level flag was true. Now stores both true and false values, consistent with preserveOpenAIDeveloperRole handling.

Made-with: Cursor
2026-03-21 16:38:46 +08:00
zhang-qiang e96b023d04 fix(ci): reword comment in default.ts to avoid t11 any-budget false positive
The word 'any' in a JSDoc comment was matched by the regex-based t11 checker. Reworded to 'prefixes' to eliminate the false positive.

Made-with: Cursor
2026-03-21 16:33:44 +08:00
zhang-qiang 7ac1d4621b Merge remote-tracking branch 'upstream/main' into feat/zws-v6 2026-03-21 16:32:57 +08:00
zhang-qiang a2d7cbe8fe feat(compat): per-protocol model compatibility config (V5)
Add per-protocol compatibility options (compatByProtocol) allowing users to configure normalizeToolCallId and preserveOpenAIDeveloperRole per client request protocol (OpenAI Chat, Responses API, Anthropic Messages) instead of globally. Includes frontend Map lookup optimization, type safety improvements, and client-safe constant extraction.

Made-with: Cursor
2026-03-21 15:23:42 +08:00
diegosouzapw c74ed29739 chore(release): v2.9.0 — cross-platform machineId, per-key rate limits, streaming cache, Alibaba DashScope, search analytics, ZWS v5, 8 issues closed
Build Electron Desktop App / Validate version (push) Failing after 33s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
2026-03-20 20:12:34 -03:00
diegosouzapw 6c8501f122 fix: cross-platform machineId without process.platform branching (#506)
Rewrite getMachineIdRaw() to use a try/catch waterfall instead of
process.platform conditionals. Next.js SWC bundler evaluates
process.platform at BUILD time, so when built on Linux, the win32
branch was dead-code-eliminated — causing 'head is not recognized'
errors on Windows.

New approach:
1. Try Windows REG.exe (existsSync check, not platform check)
2. Try macOS ioreg command
3. Try reading /etc/machine-id directly (no head/pipe)
4. Try hostname command
5. Fallback to os.hostname()

Also eliminates the patch-machine-id.cjs post-install workaround.
2026-03-20 20:07:19 -03:00
diegosouzapw 941e945f74 Merge branch 'feat/zws-v5' 2026-03-20 19:36:25 -03:00
diegosouzapw f2844d59e4 Merge branch 'feat/search-provider-routing' 2026-03-20 19:36:17 -03:00
diegosouzapw 047ff187f6 Merge branch 'feat/custom-endpoint-paths'
# Conflicts:
#	src/shared/constants/providers.ts
2026-03-20 19:34:10 -03:00
diegosouzapw 1136c40811 Merge branch 'fix/tools-filter-claude-format' 2026-03-20 19:33:08 -03:00
diegosouzapw 5a78dc864f Merge branch 'fix/issue-456-458-combo-schema-mitm-windows' 2026-03-20 19:33:08 -03:00
diegosouzapw 15c98c3048 Merge branch 'fix/developer-role-param-error' 2026-03-20 19:33:07 -03:00
diegosouzapw 0a5b005ce5 fix: resolve multiple issues (#493, #490, #452)
- #493: Fix custom provider model naming — removed incorrect prefix
  stripping in DefaultExecutor.transformRequest() that broke org-scoped
  model IDs like 'zai-org/GLM-5-FP8'

- #490: Enable context cache protection for streaming responses using
  TransformStream to inject omniModel tag as final SSE content delta
  before [DONE] marker

- #452: Add per-API-key request-count limits (max_requests_per_day,
  max_requests_per_minute) with in-memory sliding window counter,
  schema auto-migration, and Check 5 in enforceApiKeyPolicy()
2026-03-20 19:26:21 -03:00
diegosouzapw 4d64e64127 fix: KIRO MITM card text + v2.8.9 release (#505)
Build Electron Desktop App / Validate version (push) Failing after 38s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
2026-03-20 16:14:49 -03:00
Diego Rodrigues de Sa e Souza 5470c70cd0 Merge pull request #497 from zhangqiang8vip/feat/zws-v5
fix(perf): resolve dev-mode HMR resource leaks, Edge warnings, and Windows test stability
2026-03-20 16:13:27 -03:00
diegosouzapw 47959ee395 Merge branch 'main' into feat/zws-v5 2026-03-20 16:10:59 -03:00
Diego Rodrigues de Sa e Souza 7c34c178cd Merge pull request #503 from diegosouzapw/dependabot/github_actions/docker/login-action-4
chore(deps): bump docker/login-action from 3 to 4
2026-03-20 16:07:00 -03:00
Diego Rodrigues de Sa e Souza ac7cb41483 Merge pull request #502 from diegosouzapw/dependabot/github_actions/docker/setup-qemu-action-4
chore(deps): bump docker/setup-qemu-action from 3 to 4
2026-03-20 16:06:58 -03:00
Diego Rodrigues de Sa e Souza 0ab388b88e Merge pull request #501 from diegosouzapw/dependabot/github_actions/peter-evans/dockerhub-description-5
chore(deps): bump peter-evans/dockerhub-description from 4 to 5
2026-03-20 16:06:56 -03:00
Diego Rodrigues de Sa e Souza 54448902f1 Merge pull request #500 from diegosouzapw/dependabot/github_actions/actions/checkout-6
chore(deps): bump actions/checkout from 4 to 6
2026-03-20 16:06:53 -03:00
Diego Rodrigues de Sa e Souza 12107a02fd Merge pull request #499 from diegosouzapw/dependabot/github_actions/docker/build-push-action-7
chore(deps): bump docker/build-push-action from 6 to 7
2026-03-20 16:06:50 -03:00
Diego Rodrigues de Sa e Souza eace06efdc Merge pull request #498 from Sajid11194/fix/windows-machine-id-undefined-reg-exe
Thanks @Sajid11194 for fixing the Windows machine ID crash! Merged and will be part of v2.8.9. 🎉
2026-03-20 16:06:15 -03:00
dependabot[bot] ee0afa1eec chore(deps): bump docker/login-action from 3 to 4
Bumps [docker/login-action](https://github.com/docker/login-action) from 3 to 4.
- [Release notes](https://github.com/docker/login-action/releases)
- [Commits](https://github.com/docker/login-action/compare/v3...v4)

---
updated-dependencies:
- dependency-name: docker/login-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-03-20 18:26:04 +00:00
dependabot[bot] 83cdd0dafe chore(deps): bump docker/setup-qemu-action from 3 to 4
Bumps [docker/setup-qemu-action](https://github.com/docker/setup-qemu-action) from 3 to 4.
- [Release notes](https://github.com/docker/setup-qemu-action/releases)
- [Commits](https://github.com/docker/setup-qemu-action/compare/v3...v4)

---
updated-dependencies:
- dependency-name: docker/setup-qemu-action
  dependency-version: '4'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-03-20 18:25:58 +00:00
dependabot[bot] 5be025f1d1 chore(deps): bump peter-evans/dockerhub-description from 4 to 5
Bumps [peter-evans/dockerhub-description](https://github.com/peter-evans/dockerhub-description) from 4 to 5.
- [Release notes](https://github.com/peter-evans/dockerhub-description/releases)
- [Commits](https://github.com/peter-evans/dockerhub-description/compare/v4...v5)

---
updated-dependencies:
- dependency-name: peter-evans/dockerhub-description
  dependency-version: '5'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-03-20 18:25:55 +00:00
dependabot[bot] c651842ea1 chore(deps): bump actions/checkout from 4 to 6
Bumps [actions/checkout](https://github.com/actions/checkout) from 4 to 6.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/v4...v6)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-03-20 18:25:51 +00:00
dependabot[bot] 423abe6788 chore(deps): bump docker/build-push-action from 6 to 7
Bumps [docker/build-push-action](https://github.com/docker/build-push-action) from 6 to 7.
- [Release notes](https://github.com/docker/build-push-action/releases)
- [Commits](https://github.com/docker/build-push-action/compare/v6...v7)

---
updated-dependencies:
- dependency-name: docker/build-push-action
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-03-20 18:25:45 +00:00
diegosouzapw 4003c38fd1 fix: OAuth batch test crash + Test All button on provider pages (v2.8.8)
Build Electron Desktop App / Validate version (push) Failing after 35s
Build Electron Desktop App / Build Electron (macos-arm64) (push) Has been skipped
Build Electron Desktop App / Build Electron (linux) (push) Has been skipped
Build Electron Desktop App / Build Electron (macos-intel) (push) Has been skipped
Build Electron Desktop App / Build Electron (windows) (push) Has been skipped
Build Electron Desktop App / Create Release (push) Has been skipped
2026-03-20 15:09:48 -03:00
Sajid 3e0c322fd4 fix: address Gemini code review — use execFileSync and optional chaining
- Replace execSync template string with execFileSync + args array on Windows
  to prevent command injection via SystemRoot/windir environment variables
- Add optional chaining (?.) and nullish coalescing (?? "") on Windows
  REG_SZ output parsing to prevent crash if REG.exe output is unexpected
- Add optional chaining on macOS IOPlatformUUID parsing for the same reason

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 23:44:15 +06:00
zhang-qiang 7fcdd4abdd fix(ci): resolve t11 any-budget false positive and e2e bailian validation test
- Replace 'any other path' with 'all other paths' in translator comment to avoid false match by the \bany\b regex in check-t11-any-budget

- Scope e2e error locator to dialog and use .first() to prevent Playwright strict-mode violations from broad page-level selectors

- Fix fallback logic: treat dialog-still-open as validation success signal

Made-with: Cursor
2026-03-21 01:19:44 +08:00
zhang-qiang 3f3280b2d4 Merge remote-tracking branch 'upstream/main' into feat/zws-v5 2026-03-21 00:55:57 +08:00
zhang-qiang aae2399631 fix(perf): resolve HMR singleton leaks, Edge warnings, and test stability
- Use globalThis singleton guards for DB connection, HealthCheck timers, console interceptor, and graceful shutdown to survive Webpack HMR re-evaluation (fixes 485+ leaked DB connections per session)

- Split instrumentation.ts into instrumentation-node.ts with computed import path to prevent Turbopack Edge bundler from tracing Node.js modules (eliminates 10+ spurious warnings per hot compile)

- Parallelize startup imports in instrumentation-node.ts (3 batch Promise.all instead of 9 serial awaits)

- Add OMNIROUTE_USE_TURBOPACK=1 env switch in run-next.mjs (default behavior unchanged)

- Replace node:crypto with crypto in proxies.ts and errorResponse.ts to fix UnhandledSchemeError

- Add unlinkFileWithRetry with EBUSY/EPERM retry for Windows file handle timing in backup restore

- Fix pre-restore backup to await completion before closing DB

- Fix bootstrap-env, domain-persistence, and fixes-p1 test stability on Windows

Made-with: Cursor
2026-03-21 00:50:07 +08:00
Sajid 03bd2b6803 fix: resolve Windows machine ID failure due to node-machine-id bundle-time platform detection
Problem:
node-machine-id constructs the REG.exe command path at module load time
using process.platform. When Next.js bundles this module, process.platform
is "" (not "win32") in the webpack/build context, so the lookup returns
undefined and bakes "undefined\REG.exe ..." permanently into the compiled
chunk. At runtime on Windows this causes:

  Error: Command failed: undefined\REG.exe QUERY HKEY_LOCAL_MACHINE\...
  The system cannot find the path specified.

Fix:
Remove the node-machine-id dependency from machineId.ts and replace it
with a direct execSync implementation that resolves process.env.SystemRoot
at call time (not load time), so the correct Windows path is always used
regardless of when or how the module was bundled.

Platform support is preserved for Windows, macOS, and Linux/FreeBSD using
the same underlying OS queries that node-machine-id used internally.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 22:01:48 +06:00
diegosouzapw c54a57838e fix: cleanup PR #494 — remove ZWS_README, fix KIRO MITM card (#487), generify AntigravityToolCard 2026-03-20 12:19:33 -03:00
zhang-qiang 1a099ea2f2 feat(zws-v2): model compat, provider-models hardening, provider page types
- roleNormalizer/translator: ZWS v2 role handling and comments

- models + schemas: compat overrides, nullable preserveOpenAIDeveloperRole

- provider-models API: generic GET 500; compatOnly validates known provider

- providers [id] page: typed props; minimal saveModelCompatFlags PATCH

Made-with: Cursor
2026-03-20 23:03:52 +08:00
zhang-qiang 13c45807ef feat: protocol-scoped model compat (V3)
- compatByProtocol per openai/openai-responses/claude

- getters take sourceFormat; chatCore passes it

- UI: protocol selector in compat popover, dark mode select

- shared/constants/modelCompat for client-safe import (fix node:crypto build)

- ZWS_README_V3.md

Made-with: Cursor
2026-03-20 22:06:03 +08:00
diegosouzapw 00df10c29a "fix: resolved UI combo setting schema strip (#458)"
"fix: safe crypto fallback for MITM on windows (#456)"
2026-03-18 17:16:30 -03:00
diegosouzapw 41d91d628a feat(search/analytics): add Search tab to analytics dashboard + GET /api/v1/search/analytics
- SearchAnalyticsTab: provider breakdown, cache hit rate, cost summary, KPI cards
- /api/v1/search/analytics: query call_logs (request_type='search') for stats
- analytics/page.tsx: added 'Search' tab alongside Overview and Evals

Closes missing dashboard tracking identified in PR review.
2026-03-17 16:15:28 -03:00
diegosouzapw 605c3f9be1 feat(provider): add Alibaba Cloud DashScope + path validation for custom endpoint paths
- Add Alibaba Cloud (DashScope) as OpenAI-compatible provider with 12 Qwen models:
  qwen-max, qwen-plus, qwen-turbo, qwen3-coder-plus/flash, qwq-plus,
  qwq-32b, qwen3-32b, qwen3-235b-a22b
  International endpoint: dashscope-intl.aliyuncs.com/compatible-mode/v1
  Auth: Bearer API key (same as groq/xai/mistral)

- Add path traversal protection to custom endpoint paths (PR #400):
  sanitizePath() validates chatPath/modelsPath values:
  must start with '/', no '..' segments, no null bytes, max 512 chars

Closes #400 (custom endpoint paths), part of Alibaba provider integration
2026-03-16 09:44:17 -03:00
diegosouzapw 2f0894c220 test: add unit tests for Anthropic-format tools filter fix (PR #397)
8 tests covering:
- Valid OpenAI format tools (tool.function.name) preserved
- Valid Anthropic format tools (tool.name) preserved
- Empty names in both formats filtered
- Mixed format array handling
- Null/whitespace edge cases

Regression tests verify the fix from PR #397 prevents all anthropic-
format tools from being silently dropped by the empty-name filter.
2026-03-16 09:38:34 -03:00
139 changed files with 14859 additions and 1281 deletions
+108 -20
View File
@@ -4,16 +4,36 @@ description: Create a new release, bump version up to 1.x.10 threshold, update c
# Generate Release Workflow
Bump version, finalize CHANGELOG, commit, tag, push, publish to npm, and create GitHub release.
Bump version, finalize CHANGELOG, commit, open a **PR to main** and wait for user confirmation before tagging, publishing, and deploying.
> **VERSION RULE: Always use PATCH bumps (2.x.y → 2.x.y+1)**
> NEVER use `npm version minor` or `npm version major`.
> Always use: `npm version patch --no-git-tag-version`
> The threshold rule: when `y` reaches 10, bump to `2.(x+1).0` — e.g. `2.1.10` → `2.2.0`.
## Steps
---
### 1. Determine new version
## ⚠️ Two-Phase Flow
```
Phase 1 (automated): bump → docs → i18n → commit → push → open PR
↕ 🛑 STOP: Notify user, wait for PR confirmation
Phase 2 (post-merge): tag → publish → GitHub release → Docker → deploy
```
**NEVER push directly to main or create tags before the user confirms the PR.**
---
## Phase 1: Pre-Merge
### 1. Create release branch
```bash
git checkout -b release/v2.x.y
```
### 2. Determine new version
Check current version in `package.json` and increment the **patch** number only:
@@ -27,11 +47,6 @@ Version format: `2.x.y` — examples:
- `2.1.9``2.1.10` (patch)
- `2.1.10``2.2.0` (minor threshold — do manually with `sed`)
```bash
# ALWAYS use patch:
npm version patch --no-git-tag-version
```
> **⚠️ ATOMIC COMMIT RULE — Version bump MUST happen before committing feature files.**
>
> **CORRECT order:**
@@ -53,7 +68,7 @@ npm version patch --no-git-tag-version
> This ensures that `git show v2.x.y` always contains both code changes and the version bump together.
> The GitHub release tag will point to a commit that includes ALL changes for that version.
### 2. Regenerate lock file (REQUIRED after version bump)
### 3. Regenerate lock file (REQUIRED after version bump)
**Mandatory** — skipping causes `@swc/helpers` lock mismatch and CI failures:
@@ -61,7 +76,7 @@ npm version patch --no-git-tag-version
npm install
```
### 3. Finalize CHANGELOG.md
### 4. Finalize CHANGELOG.md
Replace `[Unreleased]` header with the new version and date.
Keep an empty `## [Unreleased]` section above it.
@@ -74,7 +89,7 @@ Keep an empty `## [Unreleased]` section above it.
## [2.x.y] — YYYY-MM-DD
```
### 4. Update openapi.yaml version ⚠️ MANDATORY
### 5. Update openapi.yaml version ⚠️ MANDATORY
> **CI will fail** if `docs/openapi.yaml` version ≠ `package.json` version (`check:docs-sync` enforces this).
@@ -84,33 +99,97 @@ Keep an empty `## [Unreleased]` section above it.
VERSION=$(node -p "require('./package.json').version") && sed -i "s/ version: .*/ version: $VERSION/" docs/openapi.yaml && echo "✓ openapi.yaml → $VERSION"
```
### 5. Stage, commit, and tag
### 6. Update README.md and i18n docs
Run `/update-docs` workflow steps to:
- Update feature table rows in `README.md`
- Sync changes to all 29 language `docs/i18n/*/README.md` files
- Update `docs/FEATURES.md` if Settings section changed
### 7. Run tests
// turbo
```bash
npm test
```
All tests must pass before creating the PR.
### 8. Stage, commit, and push
// turbo-all
```bash
git add package.json package-lock.json CHANGELOG.md docs/openapi.yaml
git add -A
git commit -m "chore(release): v2.x.y — summary of changes"
git push origin release/v2.x.y
```
### 9. Open PR to main
```bash
gh pr create \
--repo diegosouzapw/OmniRoute \
--base main \
--head release/v2.x.y \
--title "chore(release): v2.x.y — summary" \
--body "## 🚀 Release v2.x.y
### Changes
...
### Tests
- X/X tests pass
### ⚠️ After merging: run Phase 2 steps to tag, publish, and deploy."
```
### 10. 🛑 STOP — Notify User & Await PR Confirmation
**This is a mandatory stop point.** Use `notify_user` with `BlockedOnUser: true`:
Inform the user:
- PR URL
- Summary of changes
- Test results
- List of files changed
**DO NOT proceed to Phase 2 until the user confirms the PR looks good and merges it.**
---
## Phase 2: Post-Merge (only after user confirms)
> Run these steps only AFTER the user has merged the PR.
### 11. Pull main and create tag
```bash
git checkout main
git pull origin main
git tag -a v2.x.y -m "Release v2.x.y"
```
### 6. Push to GitHub
### 12. Push tag to GitHub
```bash
git push origin main --tags
git push origin --tags
```
### 7. Create GitHub release
### 13. Create GitHub release
```bash
gh release create v2.x.y --title "v2.x.y — summary" --notes "..."
```
### 8. 🐳 Trigger Docker Hub build (MANDATORY — keep npm and Docker in sync)
### 14. 🐳 Trigger Docker Hub build (MANDATORY — keep npm and Docker in sync)
> **CRITICAL**: Docker Hub and npm MUST always publish the same version.
> The Docker image is built automatically via GitHub Actions when a new tag is pushed.
> After pushing the tag in step 5-6, **verify the workflow runs**:
> After pushing the tag in step 11-12, **verify the workflow runs**:
```bash
# Verify the Docker workflow triggered
@@ -129,7 +208,7 @@ If the Docker build was not triggered automatically, trigger it manually:
gh workflow run docker-publish.yml --repo diegosouzapw/OmniRoute --ref v2.x.y
```
### 9. Deploy to BOTH VPS environments (MANDATORY)
### 15. Deploy to BOTH VPS environments (MANDATORY)
> Always deploy to **both** environments after every release.
> See `/deploy-vps` workflow for detailed steps.
@@ -151,18 +230,27 @@ curl -s -o /dev/null -w "LOCAL: HTTP %{http_code}\n" http://192.168.0.15:20128/
curl -s -o /dev/null -w "AKAMAI: HTTP %{http_code}\n" http://69.164.221.35:20128/
```
### 16. Clean up release branch
```bash
git branch -d release/v2.x.y
```
---
## Notes
- Always run `/update-docs` BEFORE this workflow (ensures CHANGELOG and README are current)
- The `prepublishOnly` script runs `npm run build:cli` automatically during `npm publish`
- After npm publish, verify with `npm info omniroute version`
- Lock file sync errors are caused by skipping `npm install` after version bump
- Use `gh auth switch -u diegosouzapw` if git push fails with wrong account
## Known CI Pitfalls
| CI failure | Cause | Fix |
| ------------------------------------------------------------------------- | -------------------------------------------------------- | ---------------------------------------------------------------------- |
| `[docs-sync] FAIL - OpenAPI version differs from package.json` | Skipped step 4`docs/openapi.yaml` version not updated | Run step 4 (`sed -i ...`) and commit |
| `[docs-sync] FAIL - OpenAPI version differs from package.json` | Skipped step 5`docs/openapi.yaml` version not updated | Run step 5 (`sed -i ...`) and commit |
| `[docs-sync] FAIL - CHANGELOG.md first section must be "## [Unreleased]"` | `## [Unreleased]` missing or not at top of CHANGELOG | Add `## [Unreleased]\n\n---\n` before the first versioned `## [x.y.z]` |
| Electron Linux `.deb` build fails (`FpmTarget` error) | `fpm` Ruby gem not installed on `ubuntu-latest` runner | Already fixed in `electron-release.yml` (`gem install fpm` step) |
| Docker Hub `502 error writing layer blob` | Transient Docker Hub network error during ARM64 push | Re-run the Docker publish workflow; no code change needed |
+5 -5
View File
@@ -21,18 +21,18 @@ jobs:
IMAGE_NAME: diegosouzapw/omniroute
steps:
- name: Checkout
uses: actions/checkout@v4
uses: actions/checkout@v6
with:
ref: ${{ github.event_name == 'workflow_dispatch' && format('refs/tags/v{0}', inputs.version) || '' }}
- name: Set up QEMU (for multi-arch builds)
uses: docker/setup-qemu-action@v3
uses: docker/setup-qemu-action@v4
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Login to Docker Hub
uses: docker/login-action@v3
uses: docker/login-action@v4
with:
username: ${{ secrets.DOCKERHUB_USERNAME }}
password: ${{ secrets.DOCKERHUB_TOKEN }}
@@ -50,7 +50,7 @@ jobs:
echo "Publishing Docker image: $IMAGE_NAME:$VERSION"
- name: Build and push multi-arch image
uses: docker/build-push-action@v6
uses: docker/build-push-action@v7
with:
context: .
target: runner-base
@@ -70,7 +70,7 @@ jobs:
docker buildx imagetools inspect "${{ env.IMAGE_NAME }}:${{ steps.version.outputs.version }}"
- name: Update Docker Hub description
uses: peter-evans/dockerhub-description@v4
uses: peter-evans/dockerhub-description@v5
with:
username: ${{ secrets.DOCKERHUB_USERNAME }}
password: ${{ secrets.DOCKERHUB_TOKEN }}
+1
View File
@@ -89,6 +89,7 @@ docs/*
!docs/MCP-SERVER.md
!docs/CLI-TOOLS.md
# open-sse tests
open-sse/test/*
+405
View File
@@ -2,6 +2,411 @@
## [Unreleased]
> **Coming next** — see [3.0.0-rc branch](https://github.com/diegosouzapw/OmniRoute/tree/3.0.0-rc).
---
## [3.0.0-rc.7] — 2026-03-23 _(What's New vs v2.9.5 — will be released as v3.0.0)_
> **Upgrade from v2.9.5:** 16 issues resolved · 2 community PRs merged · 2 new providers · 7 new API endpoints · 3 new features · DB migration 008+009 · 832 tests passing · 15 sub2api gap improvements (T01T15 complete).
### 🆕 New Providers
| Provider | Alias | Tier | Notes |
| ---------------- | -------------- | ---- | -------------------------------------------------------------- |
| **OpenCode Zen** | `opencode-zen` | Free | 3 models via `opencode.ai/zen/v1` (PR #530 by @kang-heewon) |
| **OpenCode Go** | `opencode-go` | Paid | 4 models via `opencode.ai/zen/go/v1` (PR #530 by @kang-heewon) |
Both providers use the new `OpencodeExecutor` with multi-format routing (`/chat/completions`, `/messages`, `/responses`, `/models/{model}:generateContent`).
---
### ✨ New Features
#### 🔑 Registered Keys Provisioning API (#464)
Auto-generate and issue OmniRoute API keys programmatically with per-provider and per-account quota enforcement.
| Endpoint | Method | Description |
| ------------------------------------- | --------- | ------------------------------------------------ |
| `/api/v1/registered-keys` | `POST` | Issue a new key — raw key returned **once only** |
| `/api/v1/registered-keys` | `GET` | List registered keys (masked) |
| `/api/v1/registered-keys/{id}` | `GET` | Get key metadata |
| `/api/v1/registered-keys/{id}` | `DELETE` | Revoke a key |
| `/api/v1/registered-keys/{id}/revoke` | `POST` | Revoke (for clients without DELETE support) |
| `/api/v1/quotas/check` | `GET` | Pre-validate quota before issuing |
| `/api/v1/providers/{id}/limits` | `GET/PUT` | Configure per-provider issuance limits |
| `/api/v1/accounts/{id}/limits` | `GET/PUT` | Configure per-account issuance limits |
| `/api/v1/issues/report` | `POST` | Report quota events to GitHub Issues |
**DB — Migration 008:** Three new tables: `registered_keys`, `provider_key_limits`, `account_key_limits`.
**Security:** Keys stored as SHA-256 hashes. Raw key shown once on creation, never retrievable again.
**Quota types:** `maxActiveKeys`, `dailyIssueLimit`, `hourlyIssueLimit` per provider and per account.
**Idempotency:** `idempotency_key` field prevents duplicate issuance. Returns `409 IDEMPOTENCY_CONFLICT` if key was already used.
**Budget per key:** `dailyBudget` / `hourlyBudget` — limits how many requests a key can route per window.
**GitHub reporting:** Optional. Set `GITHUB_ISSUES_REPO` + `GITHUB_ISSUES_TOKEN` to auto-create GitHub issues on quota exceeded or issuance failures.
#### 🎨 Provider Icons — @lobehub/icons (#529)
All provider icons in the dashboard now use `@lobehub/icons` React components (130+ providers with SVG).
Fallback chain: **Lobehub SVG → existing `/providers/{id}.png` → generic icon**. Uses a proper React `ErrorBoundary` pattern.
#### 🔄 Model Auto-Sync Scheduler (#488)
OmniRoute now automatically refreshes model lists for connected providers every **24 hours**.
- Runs on server startup via the existing `/api/sync/initialize` hook
- Configurable via `MODEL_SYNC_INTERVAL_HOURS` environment variable
- Covers 16 major providers
- Records last sync time in the settings database
---
### 🔧 Bug Fixes
#### OAuth & Auth
- **#537 — Gemini CLI OAuth:** Clear actionable error when `GEMINI_OAUTH_CLIENT_SECRET` is missing in Docker/self-hosted deployments. Previously showed cryptic `client_secret is missing` from Google. Now provides specific `docker-compose.yml` and `~/.omniroute/.env` instructions.
#### Providers & Routing
- **#536 — LongCat AI:** Fixed `baseUrl` (`api.longcat.chat/openai`) and `authHeader` (`Authorization: Bearer`).
- **#535 — Pinned model override:** `body.model` is now correctly set to `pinnedModel` when context-cache protection is active.
- **#532 — OpenCode Go key validation:** Now uses the `zen/v1` test endpoint (`testKeyBaseUrl`) — same key works for both tiers.
#### CLI & Tools
- **#527 — Claude Code + Codex loop:** `tool_result` blocks are now converted to text instead of dropped, stopping infinite tool-result loops.
- **#524 — OpenCode config save:** Added `saveOpenCodeConfig()` handler (XDG_CONFIG_HOME aware, writes TOML).
- **#521 — Login stuck:** Login no longer freezes after skipping password setup — redirects correctly to onboarding.
- **#522 — API Manager:** Removed misleading "Copy masked key" button (replaced with a lock icon tooltip).
- **#532 — OpenCode Go config:** Guide settings handler now handles `opencode` toolId.
#### Developer Experience
- **#489 — Antigravity:** Missing `googleProjectId` returns a structured 422 error with reconnect guidance instead of a cryptic crash.
- **#510 — Windows paths:** MSYS2/Git-Bash paths (`/c/Program Files/...`) are now normalized to `C:\\Program Files\\...` automatically.
- **#492 — CLI startup:** `omniroute` CLI now detects `mise`/`nvm`-managed Node when `app/server.js` is missing and shows targeted fix instructions.
---
### 📖 Documentation Updates
- **#513** — Docker password reset: `INITIAL_PASSWORD` env var workaround documented
- **#520** — pnpm: `pnpm approve-builds better-sqlite3` step documented
---
### ✅ Issues Resolved in v3.0.0
`#464` `#488` `#489` `#492` `#510` `#513` `#520` `#521` `#522` `#524` `#527` `#529` `#532` `#535` `#536` `#537`
---
### 🔀 Community PRs Merged
| PR | Author | Summary |
| -------- | ------------ | ---------------------------------------------------------------------- |
| **#530** | @kang-heewon | OpenCode Zen + Go providers with `OpencodeExecutor` and improved tests |
---
## [3.0.0-rc.7] - 2026-03-23
### 🔧 Improvements (sub2api Gap Analysis — T05, T08, T09, T13, T14)
- **T05** — Rate-limit DB persistence: `setConnectionRateLimitUntil()`, `isConnectionRateLimited()`, `getRateLimitedConnections()` in `providers.ts`. The existing `rate_limited_until` column is now exposed as a dedicated API — OAuth token refresh must NOT touch this field to prevent rate-limit loops.
- **T08** — Per-API-key session limit: `max_sessions INTEGER DEFAULT 0` added to `api_keys` via auto-migration. `sessionManager.ts` gains `registerKeySession()`, `unregisterKeySession()`, `checkSessionLimit()`, and `getActiveSessionCountForKey()`. Callers in `chatCore.js` can enforce the limit and decrement on `req.close`.
- **T09** — Codex vs Spark rate-limit scopes: `getCodexModelScope()` and `getCodexRateLimitKey()` in `codex.ts`. Standard models (`gpt-5.x-codex`, `codex-mini`) get scope `"codex"`; spark models (`codex-spark*`) get scope `"spark"`. Rate-limit keys should be `${accountId}:${scope}` so exhausting one pool doesn't block the other.
- **T13** — Stale quota display fix: `getEffectiveQuotaUsage(used, resetAt)` returns `0` when the reset window has passed; `formatResetCountdown(resetAt)` returns a human-readable countdown string (e.g. `"2h 35m"`). Both exported from `providers.ts` + `localDb.ts` for dashboard consumption.
- **T14** — Proxy fast-fail: new `src/lib/proxyHealth.ts` with `isProxyReachable(proxyUrl, timeoutMs=2000)` (TCP check, ≤2s instead of 30s timeout), `getCachedProxyHealth()`, `invalidateProxyHealth()`, and `getAllProxyHealthStatuses()`. Results cached 30s by default; configurable via `PROXY_FAST_FAIL_TIMEOUT_MS` / `PROXY_HEALTH_CACHE_TTL_MS`.
### 🧪 Tests
- Test suite: **832 tests, 0 failures**
---
## [3.0.0-rc.6] - 2026-03-23
### 🔧 Bug Fixes & Improvements (sub2api Gap Analysis — T01T15)
- **T01** — `requested_model` column in `call_logs` (migration 009): track which model the client originally requested vs the actual routed model. Enables fallback rate analytics.
- **T02** — Strip empty text blocks from nested `tool_result.content`: prevents Anthropic 400 errors (`text content blocks must be non-empty`) when Claude Code chains tool results.
- **T03** — Parse `x-codex-5h-*` / `x-codex-7d-*` headers: `parseCodexQuotaHeaders()` + `getCodexResetTime()` extract Codex quota windows for precise cooldown scheduling instead of generic 5-min fallback.
- **T04** — `X-Session-Id` header for external sticky routing: `extractExternalSessionId()` in `sessionManager.ts` reads `x-session-id` / `x-omniroute-session` headers with `ext:` prefix to avoid collision with internal SHA-256 session IDs. Nginx-compatible (hyphenated header).
- **T06** — Account deactivated → permanent block: `isAccountDeactivated()` in `accountFallback.ts` detects 401 deactivation signals and applies a 1-year cooldown to prevent retrying permanently dead accounts.
- **T07** — X-Forwarded-For IP validation: new `src/lib/ipUtils.ts` with `extractClientIp()` and `getClientIpFromRequest()` — skips `unknown`/non-IP entries in `X-Forwarded-For` chains (Nginx/proxy-forwarded requests).
- **T10** — Credits exhausted → distinct fallback: `isCreditsExhausted()` in `accountFallback.ts` returns 1h cooldown with `creditsExhausted` flag, distinct from generic 429 rate limiting.
- **T11** — `max` reasoning effort → 131072 budget tokens: `EFFORT_BUDGETS` and `THINKING_LEVEL_MAP` updated; reverse mapping now returns `"max"` for full-budget responses. Unit test updated.
- **T12** — MiniMax M2.7 pricing entries added: `minimax-m2.7`, `MiniMax-M2.7`, `minimax-m2.7-highspeed` added to pricing table (sub2api PR #1120). M2.5/GLM-4.7/GLM-5/Kimi pricing already existed.
- **T15** — Array content normalization: `normalizeContentToString()` helper in `openai-to-claude.ts` correctly collapses array-formatted system/tool messages to string before sending to Anthropic.
### 🧪 Tests
- Test suite: **832 tests, 0 failures** (unchanged from rc.5)
---
## [3.0.0-rc.5] - 2026-03-22
### ✨ New Features
- **#464** — Registered Keys Provisioning API: auto-issue API keys with per-provider & per-account quota enforcement
- `POST /api/v1/registered-keys` — issue keys with idempotency support
- `GET /api/v1/registered-keys` — list (masked) registered keys
- `GET /api/v1/registered-keys/{id}` — get key metadata
- `DELETE /api/v1/registered-keys/{id}` / `POST ../{id}/revoke` — revoke keys
- `GET /api/v1/quotas/check` — pre-validate before issuing
- `PUT /api/v1/providers/{id}/limits` — set provider issuance limits
- `PUT /api/v1/accounts/{id}/limits` — set account issuance limits
- `POST /api/v1/issues/report` — optional GitHub issue reporting
- DB migration 008: `registered_keys`, `provider_key_limits`, `account_key_limits` tables
---
## [3.0.0-rc.4] - 2026-03-22
### ✨ New Features
- **#530 (PR)** — OpenCode Zen and OpenCode Go providers added (by @kang-heewon)
- New `OpencodeExecutor` with multi-format routing (`/chat/completions`, `/messages`, `/responses`)
- 7 models across both tiers
---
## [3.0.0-rc.3] - 2026-03-22
### ✨ New Features
- **#529** — Provider icons now use [@lobehub/icons](https://github.com/lobehub/lobe-icons) with graceful PNG fallback and a `ProviderIcon` component (130+ providers supported)
- **#488** — Auto-update model lists every 24h via `modelSyncScheduler` (configurable via `MODEL_SYNC_INTERVAL_HOURS`)
### 🔧 Bug Fixes
- **#537** — Gemini CLI OAuth: now shows clear actionable error when `GEMINI_OAUTH_CLIENT_SECRET` is missing in Docker/self-hosted deployments
---
## [3.0.0-rc.2] - 2026-03-22
### 🔧 Bug Fixes
- **#536** — LongCat AI key validation: fixed baseUrl (`api.longcat.chat/openai`) and authHeader (`Authorization: Bearer`)
- **#535** — Pinned model override: `body.model` is now set to `pinnedModel` when context-cache protection detects a pinned model
- **#524** — OpenCode config now saved correctly: added `saveOpenCodeConfig()` handler (XDG_CONFIG_HOME aware, writes TOML)
---
## [3.0.0-rc.1] - 2026-03-22
### 🔧 Bug Fixes
- **#521** — Login no longer gets stuck after skipping password setup (redirects to onboarding)
- **#522** — API Manager: Removed misleading "Copy masked key" button (replaced with lock icon tooltip)
- **#527** — Claude Code + Codex superpowers loop: `tool_result` blocks now converted to text instead of dropped
- **#532** — OpenCode GO API key validation now uses the correct `zen/v1` endpoint (`testKeyBaseUrl`)
- **#489** — Antigravity: missing `googleProjectId` returns structured 422 error with reconnect guidance
- **#510** — Windows: MSYS2/Git-Bash paths (`/c/Program Files/...`) are now normalized to `C:\\Program Files\\...`
- **#492** — `omniroute` CLI now detects `mise`/`nvm` when `app/server.js` is missing and shows targeted fix
### 📖 Documentation
- **#513** — Docker password reset: `INITIAL_PASSWORD` env var workaround documented
- **#520** — pnpm: `pnpm approve-builds better-sqlite3` documented
### ✅ Closed Issues
#489, #492, #510, #513, #520, #521, #522, #525, #527, #532
---
## [2.9.5] — 2026-03-22
> Sprint: New OpenCode providers, embedding credentials fix, CLI masked key bug, CACHE_TAG_PATTERN fix.
### 🐛 Bug Fixes
- **CLI tools save masked API key to config files** — `claude-settings`, `cline-settings`, and `openclaw-settings` POST routes now accept a `keyId` param and resolve the real API key from DB before writing to disk. `ClaudeToolCard` updated to send `keyId` instead of the masked display string. Fixes #523, #526.
- **Custom embedding providers: `No credentials` error** — `/v1/embeddings` now tracks `credentialsProviderId` separately from the routing prefix, so credentials are fetched from the matching provider node ID rather than the public prefix string. Fixes a regression where `google/gemini-embedding-001` and similar custom-provider models would always fail with a credentials error. Fixes #532-related. (PR #528 by @jacob2826)
- **Context cache protection regex misses `\n` prefix** — `CACHE_TAG_PATTERN` in `comboAgentMiddleware.ts` updated to match both literal `\n` (backslash-n) and actual newline U+000A that `combo.ts` streaming injects around the `<omniModel>` tag after fix #515. Fixes #531.
### ✨ New Providers
- **OpenCode Zen** — Free tier gateway at `opencode.ai/zen/v1` with 3 models: `minimax-m2.5-free`, `big-pickle`, `gpt-5-nano`
- **OpenCode Go** — Subscription service at `opencode.ai/zen/go/v1` with 4 models: `glm-5`, `kimi-k2.5`, `minimax-m2.7` (Claude format), `minimax-m2.5` (Claude format)
- Both providers use the new `OpencodeExecutor` which routes dynamically to `/chat/completions`, `/messages`, `/responses`, or `/models/{model}:generateContent` based on the requested model. (PR #530 by @kang-heewon)
---
## [2.9.4] — 2026-03-21
> Sprint: Bug fixes — preserve Codex prompt cache key, fix tagContent JSON escaping, sync expired token status to DB.
### 🐛 Bug Fixes
- **fix(translator)**: Preserve `prompt_cache_key` in Responses API → Chat Completions translation (#517)
— The field is a cache-affinity signal used by Codex; stripping it was preventing prompt cache hits.
Fixed in `openai-responses.ts` and `responsesApiHelper.ts`.
- **fix(combo)**: Escape `\n` in `tagContent` so injected JSON string is valid (#515)
— Template literal newlines (U+000A) are not allowed unescaped inside JSON string values.
Replaced with `\\n` literal sequences in `open-sse/services/combo.ts`.
- **fix(usage)**: Sync expired token status back to DB on live auth failure (#491)
— When the Limits & Quotas live check returns 401/403, the connection `testStatus` is now updated
to `"expired"` in the database so the Providers page reflects the same degraded state.
Fixed in `src/app/api/usage/[connectionId]/route.ts`.
---
## [2.9.3] — 2026-03-21
> Sprint: Add 5 new free AI providers — LongCat, Pollinations, Cloudflare AI, Scaleway, AI/ML API.
### ✨ New Providers
- **feat(providers/longcat)**: Add LongCat AI (`lc/`) — 50M tokens/day free (Flash-Lite) + 500K/day (Chat/Thinking) during public beta. OpenAI-compatible, standard Bearer auth.
- **feat(providers/pollinations)**: Add Pollinations AI (`pol/`) — no API key required. Proxies GPT-5, Claude, Gemini, DeepSeek V3, Llama 4 (1 req/15s free). Custom executor handles optional auth.
- **feat(providers/cloudflare-ai)**: Add Cloudflare Workers AI (`cf/`) — 10K Neurons/day free (~150 LLM responses or 500s Whisper audio). 50+ models on global edge. Custom executor builds dynamic URL with `accountId` from credentials.
- **feat(providers/scaleway)**: Add Scaleway Generative APIs (`scw/`) — 1M free tokens for new accounts. EU/GDPR compliant (Paris). Qwen3 235B, Llama 3.1 70B, Mistral Small 3.2.
- **feat(providers/aimlapi)**: Add AI/ML API (`aiml/`) — $0.025/day free credit, 200+ models (GPT-4o, Claude, Gemini, Llama) via single aggregator endpoint.
### 🔄 Provider Updates
- **feat(providers/together)**: Add `hasFree: true` + 3 permanently free model IDs: `Llama-3.3-70B-Instruct-Turbo-Free`, `Llama-Vision-Free`, `DeepSeek-R1-Distill-Llama-70B-Free`
- **feat(providers/gemini)**: Add `hasFree: true` + `freeNote` (1,500 req/day, no credit card needed, aistudio.google.com)
- **chore(providers/gemini)**: Rename display name to `Gemini (Google AI Studio)` for clarity
### ⚙️ Infrastructure
- **feat(executors/pollinations)**: New `PollinationsExecutor` — omits `Authorization` header when no API key provided
- **feat(executors/cloudflare-ai)**: New `CloudflareAIExecutor` — dynamic URL construction requires `accountId` in provider credentials
- **feat(executors)**: Register `pollinations`, `pol`, `cloudflare-ai`, `cf` executor mappings
### 📝 Documentation
- **docs(readme)**: Expanded free combo stack to 11 providers ($0 forever)
- **docs(readme)**: Added 4 new free provider sections (LongCat, Pollinations, Cloudflare AI, Scaleway) with model tables
- **docs(readme)**: Updated pricing table with 4 new free tier rows
- **docs(i18n/pt-BR)**: Updated pricing table + added LongCat/Pollinations/Cloudflare AI/Scaleway sections in Portuguese
- **docs(new-features/ai)**: 10 task spec files + master implementation plan in `docs/new-features/ai/`
### 🧪 Tests
- Test suite: **821 tests, 0 failures** (unchanged)
---
## [2.9.2] — 2026-03-21
> Sprint: Fix media transcription (Deepgram/HuggingFace Content-Type, language detection) and TTS error display.
### 🐛 Bug Fixes
- **fix(transcription)**: Deepgram and HuggingFace audio transcription now correctly map `video/mp4``audio/mp4` and other media MIME types via new `resolveAudioContentType()` helper. Previously, uploading `.mp4` files consistently returned "No speech detected" because Deepgram was receiving `Content-Type: video/mp4`.
- **fix(transcription)**: Added `detect_language=true` to Deepgram requests — auto-detects audio language (Portuguese, Spanish, etc.) instead of defaulting to English. Fixes non-English transcriptions returning empty or garbage results.
- **fix(transcription)**: Added `punctuate=true` to Deepgram requests for higher-quality transcription output with correct punctuation.
- **fix(tts)**: `[object Object]` error display in Text-to-Speech responses fixed in both `audioSpeech.ts` and `audioTranscription.ts`. The `upstreamErrorResponse()` function now correctly extracts nested string messages from providers like ElevenLabs that return `{ error: { message: "...", status_code: 401 } }` instead of a flat error string.
### 🧪 Tests
- Test suite: **821 tests, 0 failures** (unchanged)
### Triaged Issues
- **#508** — Tool call format regression: requested proxy logs and provider chain info (`needs-info`)
- **#510** — Windows CLI healthcheck path: requested shell/Node version info (`needs-info`)
- **#485** — Kiro MCP tool calls: closed as external Kiro issue (not OmniRoute)
- **#442** — Baseten /models endpoint: closed (documented manual workaround)
- **#464** — Key provisioning API: acknowledged as roadmap item
---
## [2.9.1] — 2026-03-21
> Sprint: Fix SSE omniModel data loss, merge per-protocol model compatibility.
### Bug Fixes
- **#511** — Critical: `<omniModel>` tag was sent after `finish_reason:stop` in SSE streams, causing data loss. Tag is now injected into the first non-empty content chunk, guaranteeing delivery before SDKs close the connection.
### Merged PRs
- **PR #512** (@zhangqiang8vip): Per-protocol model compatibility — `normalizeToolCallId` and `preserveOpenAIDeveloperRole` can now be configured per client protocol (OpenAI, Claude, Responses API). New `compatByProtocol` field in model config with Zod validation.
### Triaged Issues
- **#510** — Windows CLI healthcheck_failed: requested PATH/version info
- **#509** — Turbopack Electron regression: upstream Next.js bug, documented workarounds
- **#508** — macOS black screen: suggested `--disable-gpu` workaround
---
## [2.9.0] — 2026-03-20
> Sprint: Cross-platform machineId fix, per-API-key rate limits, streaming context cache, Alibaba DashScope, search analytics, ZWS v5, and 8 issues closed.
### ✨ New Features
- **feat(search)**: Search Analytics tab in `/dashboard/analytics` — provider breakdown, cache hit rate, cost tracking. New API: `GET /api/v1/search/analytics` (#feat/search-provider-routing)
- **feat(provider)**: Alibaba Cloud DashScope added with custom endpoint path validation — configurable `chatPath` and `modelsPath` per node (#feat/custom-endpoint-paths)
- **feat(api)**: Per-API-key request-count limits — `max_requests_per_day` and `max_requests_per_minute` columns with in-memory sliding-window enforcement returning HTTP 429 (#452)
- **feat(dev)**: ZWS v5 — HMR leak fix (485 DB connections → 1), memory 2.4GB → 195MB, `globalThis` singletons, Edge Runtime warning fix (@zhangqiang8vip)
### 🐛 Bug Fixes
- **fix(#506)**: Cross-platform `machineId``getMachineIdRaw()` rewritten with try/catch waterfall (Windows REG.exe → macOS ioreg → Linux file read → hostname → `os.hostname()`). Eliminates `process.platform` branching that Next.js bundler dead-code-eliminated, fixing `'head' is not recognized` on Windows. Also fixes #466.
- **fix(#493)**: Custom provider model naming — removed incorrect prefix stripping in `DefaultExecutor.transformRequest()` that mangled org-scoped model IDs like `zai-org/GLM-5-FP8`.
- **fix(#490)**: Streaming + context cache protection — `TransformStream` intercepts SSE to inject `<omniModel>` tag before `[DONE]` marker, enabling context cache protection for streaming responses.
- **fix(#458)**: Combo schema validation — `system_message`, `tool_filter_regex`, `context_cache_protection` fields now pass Zod validation on save.
- **fix(#487)**: KIRO MITM card cleanup — removed ZWS_README, generified `AntigravityToolCard` to use dynamic tool metadata.
### 🧪 Tests
- Added Anthropic-format tools filter unit tests (PR #397) — 8 regression tests for `tool.name` without `.function` wrapper
- Test suite: **821 tests, 0 failures** (up from 813)
### 📋 Issues Closed (8)
- **#506** — Windows machineId `head` not recognized (fixed)
- **#493** — Custom provider model naming (fixed)
- **#490** — Streaming context cache (fixed)
- **#452** — Per-API-key request limits (implemented)
- **#466** — Windows login failure (same root cause as #506)
- **#504** — MITM inactive (expected behavior)
- **#462** — Gemini CLI PSA (resolved)
- **#434** — Electron app crash (duplicate of #402)
## [2.8.9] — 2026-03-20
> Sprint: Merge community PRs, fix KIRO MITM card, dependency updates.
### Merged PRs
- **PR #498** (@Sajid11194): Fix Windows machine ID crash (`undefined\REG.exe`). Replaces `node-machine-id` with native OS registry queries. **Closes #486.**
- **PR #497** (@zhangqiang8vip): Fix dev-mode HMR resource leaks — 485 leaked DB connections → 1, memory 2.4GB → 195MB. `globalThis` singletons, Edge Runtime warning fix, Windows test stability. (+1168/-338 across 22 files)
- **PRs #499-503** (Dependabot): GitHub Actions updates — `docker/build-push-action@7`, `actions/checkout@6`, `peter-evans/dockerhub-description@5`, `docker/setup-qemu-action@4`, `docker/login-action@4`.
### Bug Fixes
- **#505** — KIRO MITM card now displays tool-specific instructions (`api.anthropic.com`) instead of Antigravity-specific text.
- **#504** — Responded with UX clarification (MITM "Inactive" is expected behavior when proxy is not running).
---
## [2.8.8] — 2026-03-20
> Sprint: Fix OAuth batch test crash, add "Test All" button to individual provider pages.
### Bug Fixes
- **OAuth batch test crash** (ERR_CONNECTION_REFUSED): Replaced sequential for-loop with 5-connection concurrency limit + 30s per-connection timeout via `Promise.race()` + `Promise.allSettled()`. Prevents server crash when testing large OAuth provider groups (~30+ connections).
### Features
- **"Test All" button on provider pages**: Individual provider pages (e.g., `/providers/codex`) now show a "Test All" button in the Connections header when there are 2+ connections. Uses `POST /api/providers/test-batch` with `{mode: "provider", providerId}`. Results displayed in a modal with pass/fail summary and per-connection diagnosis.
---
## [2.8.7] — 2026-03-20
+106 -28
View File
@@ -11,7 +11,9 @@ _Your universal API proxy — one endpoint, 44+ providers, zero downtime. Now wi
<div align="center">
[![npm version](https://img.shields.io/npm/v/omniroute?color=cb3837&logo=npm)](https://www.npmjs.com/package/omniroute)
[![npm downloads](https://img.shields.io/npm/dm/omniroute?color=cb3837&logo=npm&label=npm%20downloads)](https://www.npmjs.com/package/omniroute)
[![Docker Hub](https://img.shields.io/docker/v/diegosouzapw/omniroute?label=Docker%20Hub&logo=docker&color=2496ED)](https://hub.docker.com/r/diegosouzapw/omniroute)
[![Docker Pulls](https://img.shields.io/docker/pulls/diegosouzapw/omniroute?logo=docker&color=2496ED&label=docker%20pulls)](https://hub.docker.com/r/diegosouzapw/omniroute)
[![License](https://img.shields.io/github/license/diegosouzapw/OmniRoute)](https://github.com/diegosouzapw/OmniRoute/blob/main/LICENSE)
[![Website](https://img.shields.io/badge/Website-omniroute.online-blue?logo=google-chrome&logoColor=white)](https://omniroute.online)
[![WhatsApp](https://img.shields.io/badge/WhatsApp-Community-25D366?logo=whatsapp&logoColor=white)](https://chat.whatsapp.com/JI7cDQ1GyaiDHhVBpLxf8b?mode=gi_t)
@@ -24,6 +26,25 @@ _Your universal API proxy — one endpoint, 44+ providers, zero downtime. Now wi
---
## 🆕 What's New in v3.0.0
> **Upgrading from v2.9.5?** — See the [full CHANGELOG](CHANGELOG.md#300--2026-03-22-release-candidate--not-yet-merged-to-main) for all changes.
| Area | Change |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🔑 **Registered Keys API** | Auto-provision API keys via `POST /api/v1/registered-keys` with per-provider/account quota enforcement, idempotency, SHA-256 storage, and optional GitHub issue reporting |
| 🎨 **Provider Icons** | 130+ provider logos via `@lobehub/icons` (SVG) with PNG → generic fallback chain |
| 🔄 **Model Auto-Sync** | 24h scheduler refreshes model lists for 16 providers on startup — configurable via `MODEL_SYNC_INTERVAL_HOURS` |
| 🌐 **OpenCode Zen/Go** | Two new providers from @kang-heewon via PR #530: free tier + subscription tier via `OpencodeExecutor` |
| 🐛 **Gemini CLI OAuth** | Actionable error when `GEMINI_OAUTH_CLIENT_SECRET` is missing in Docker (was cryptic Google error) |
| 🐛 **OpenCode config** | `saveOpenCodeConfig()` now correctly writes TOML to `XDG_CONFIG_HOME` |
| 🐛 **Pinned model override** | `body.model` correctly set to `pinnedModel` on context-cache protection |
| 🐛 **Codex/Claude loop** | `tool_result` blocks now converted to text to stop infinite loops |
| 🐛 **Login redirect** | Login no longer freezes after skipping password setup |
| 🐛 **Windows paths** | MSYS2/Git-Bash paths (`/c/...`) normalized to `C:\...` automatically |
---
## 🖼️ Main Dashboard
<div align="center">
@@ -716,7 +737,7 @@ Outcome: deep fallback depth for deadline-critical workloads
**Point any IDE/CLI to:** `http://localhost:20128/v1` · API Key: `any-string` · Done.
> **Optional extra coverage (also free):** Groq API key (30 RPM free), NVIDIA NIM (40 RPM free, 70+ models), Cerebras (1M tok/day).
> **Optional extra coverage (also free):** Groq API key (30 RPM free), NVIDIA NIM (40 RPM free, 70+ models), Cerebras (1M tok/day), LongCat API key (50M tokens/day!), Cloudflare Workers AI (10K Neurons/day, 50+ models).
## ⚡ Quick Start
@@ -921,18 +942,28 @@ When minimized, OmniRoute lives in your system tray with quick actions:
| **🆓 FREE** | iFlow | **$0** | Unlimited | 5 models unlimited |
| | Qwen | **$0** | Unlimited | 4 models unlimited |
| | Kiro | **$0** | Unlimited | Claude Sonnet/Haiku (AWS Builder) |
| | LongCat Flash-Lite 🆕 | **$0** (50M tok/day 🔥) | 1 RPS | Largest free quota on Earth |
| | Pollinations AI 🆕 | **$0** (no key needed) | 1 req/15s | GPT-5, Claude, DeepSeek, Llama 4 |
| | Cloudflare Workers AI 🆕 | **$0** (10K Neurons/day) | ~150 resp/day | 50+ models, global edge |
| | Scaleway AI 🆕 | **$0** (1M tokens total) | Rate limited | EU/GDPR, Qwen3 235B, Llama 70B |
> 🆕 **New models added (Mar 2026):** Grok-4 Fast family at $0.20/$0.50/M (benchmarked at 1143ms — 30% faster than Gemini 2.5 Flash), GLM-5 via Z.AI with 128K output, MiniMax M2.5 reasoning, DeepSeek V3.2 updated pricing, Kimi K2.5 via Moonshot direct API.
**💡 $0 Combo Stack — The Complete Free Setup:**
```
Gemini CLI (180K/mo free)
→ iFlow (unlimited: kimi-k2-thinking, qwen3-coder-plus, deepseek-r1)
→ Kiro (Claude Sonnet 4.5 + Haiku — unlimited, via AWS Builder ID)
→ Qwen (4 models — unlimited)
→ Groq (14.4K req/day — ultra-fast)
→ NVIDIA NIM (70+ models — 40 RPM forever)
# 🆓 Ultimate Free Stack 2026 — 11 Providers, $0 Forever
Kiro (kr/) → Claude Sonnet/Haiku UNLIMITED
iFlow (if/) → kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED
LongCat Lite (lc/) → LongCat-Flash-Lite — 50M tokens/day 🔥
Pollinations (pol/) → GPT-5, Claude, DeepSeek, Llama 4 — no key needed
Qwen (qw/) → qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next UNLIMITED
Gemini (gemini/) → Gemini 2.5 Flash — 1,500 req/day free API key
Cloudflare AI (cf/) → Llama 70B, Gemma 3, Mistral — 10K Neurons/day
Scaleway (scw/) → Qwen3 235B, Llama 70B — 1M free tokens (EU)
Groq (groq/) → Llama/Gemma ultra-fast — 14.4K req/day
NVIDIA NIM (nvidia/) → 70+ open models — 40 RPM forever
Cerebras (cerebras/) → Llama/Qwen world-fastest — 1M tok/day
```
**Zero cost. Never stops coding.** Configure this as one OmniRoute combo and all fallbacks happen automatically — no manual switching ever.
@@ -1003,19 +1034,66 @@ Available free: `llama-3.3-70b`, `llama-3.1-8b`, `deepseek-r1-distill-llama-70b`
Available free: `llama-3.3-70b-versatile`, `gemma2-9b-it`, `mixtral-8x7b`, `whisper-large-v3`
> **💡 The Ultimate Free Stack:**
### 🔴 LONGCAT AI (Free API Key — longcat.chat) 🆕
| Model | Prefix | Daily Free Quota | Notes |
| ----------------------------- | ------ | ----------------- | ----------------------- |
| `LongCat-Flash-Lite` | `lc/` | **50M tokens** 💥 | Largest free quota ever |
| `LongCat-Flash-Chat` | `lc/` | 500K tokens | Multi-turn chat |
| `LongCat-Flash-Thinking` | `lc/` | 500K tokens | Reasoning / CoT |
| `LongCat-Flash-Thinking-2601` | `lc/` | 500K tokens | Jan 2026 version |
| `LongCat-Flash-Omni-2603` | `lc/` | 500K tokens | Multimodal |
> 100% free while in public beta. Sign up at [longcat.chat](https://longcat.chat) with email or phone. Resets daily 00:00 UTC.
### 🟢 POLLINATIONS AI (No API Key Required) 🆕
| Model | Prefix | Rate Limit | Provider Behind |
| ---------- | ------ | ---------- | ------------------ |
| `openai` | `pol/` | 1 req/15s | GPT-5 |
| `claude` | `pol/` | 1 req/15s | Anthropic Claude |
| `gemini` | `pol/` | 1 req/15s | Google Gemini |
| `deepseek` | `pol/` | 1 req/15s | DeepSeek V3 |
| `llama` | `pol/` | 1 req/15s | Meta Llama 4 Scout |
| `mistral` | `pol/` | 1 req/15s | Mistral AI |
> ✨ **Zero friction:** No signup, no API key. Add the Pollinations provider with an empty key field and it works immediately.
### 🟠 CLOUDFLARE WORKERS AI (Free API Key — cloudflare.com) 🆕
| Tier | Daily Neurons | Equivalent Usage | Notes |
| ---- | ------------- | --------------------------------------- | ----------------------- |
| Free | **10,000** | ~150 LLM resp / 500s audio / 15K embeds | Global edge, 50+ models |
Popular free models: `@cf/meta/llama-3.3-70b-instruct`, `@cf/google/gemma-3-12b-it`, `@cf/openai/whisper-large-v3-turbo` (free audio!), `@cf/qwen/qwen2.5-coder-15b-instruct`
> Requires API Token + Account ID from [dash.cloudflare.com](https://dash.cloudflare.com). Store Account ID in provider settings.
### 🟣 SCALEWAY AI (1M Free Tokens — scaleway.com) 🆕
| Tier | Free Quota | Location | Notes |
| ---- | ------------- | ------------ | ----------------------------------- |
| Free | **1M tokens** | 🇫🇷 Paris, EU | No credit card needed within limits |
Available free: `qwen3-235b-a22b-instruct-2507` (Qwen3 235B!), `llama-3.1-70b-instruct`, `mistral-small-3.2-24b-instruct-2506`, `deepseek-v3-0324`
> EU/GDPR compliant. Get API key at [console.scaleway.com](https://console.scaleway.com).
> **💡 The Ultimate Free Stack (11 Providers, $0 Forever):**
>
> ```
> Kiro (Claude, unlimited)
> iFlow (5 models, unlimited)
> → Qwen (4 models, unlimited)
> → Gemini CLI (180K/mo)
> → Cerebras (1M tok/day)
> → Groq (14.4K req/day)
> → NVIDIA NIM (40 RPM, 70+ models)
> Kiro (kr/) → Claude Sonnet/Haiku UNLIMITED
> iFlow (if/) → kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED
> LongCat Lite (lc/) → LongCat-Flash-Lite — 50M tokens/day 🔥
> Pollinations (pol/) → GPT-5, Claude, DeepSeek, Llama 4 — no key needed
> Qwen (qw/) → qwen3-coder models UNLIMITED
> Gemini (gemini/) → Gemini 2.5 Flash — 1,500 req/day free
> Cloudflare AI (cf/) → 50+ models — 10K Neurons/day
> Scaleway (scw/) → Qwen3 235B, Llama 70B — 1M free tokens (EU)
> Groq (groq/) → Llama/Gemma — 14.4K req/day ultra-fast
> NVIDIA NIM (nvidia/) → 70+ open models — 40 RPM forever
> Cerebras (cerebras/) → Llama/Qwen world-fastest — 1M tok/day
> ```
>
> Configure this as an OmniRoute combo and you'll never pay for AI again.
## 🎙️ Free Transcription Combo
@@ -1105,17 +1183,17 @@ OmniRoute v2.0 is built as an operational platform, not just a relay proxy.
### 🎵 Multi-Modal APIs
| Feature | What It Does |
| -------------------------- | ------------------------------------------------------------------------------------------------------------ |
| 🖼️ **Image Generation** | `/v1/images/generations` with cloud and local backends |
| 📐 **Embeddings** | `/v1/embeddings` for search and RAG pipelines |
| 🎤 **Audio Transcription** | `/v1/audio/transcriptions` (Whisper and additional providers) |
| 🔊 **Text-to-Speech** | `/v1/audio/speech` (multiple engines/providers) |
| 🎬 **Video Generation** | `/v1/videos/generations` (ComfyUI + SD WebUI workflows) |
| 🎵 **Music Generation** | `/v1/music/generations` (ComfyUI workflows) |
| 🛡️ **Moderations** | `/v1/moderations` safety checks |
| 🔀 **Reranking** | `/v1/rerank` for relevance scoring |
| 🔍 **Web Search** 🆕 | `/v1/search` — 5 providers (Serper, Brave, Perplexity, Exa, Tavily), 6,500+ free/month, auto-failover, cache |
| Feature | What It Does |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Image Generation** | `/v1/images/generations` with cloud and local backends |
| 📐 **Embeddings** | `/v1/embeddings` for search and RAG pipelines |
| 🎤 **Audio Transcription** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Text-to-Speech** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) with correct error messages |
| 🎬 **Video Generation** | `/v1/videos/generations` (ComfyUI + SD WebUI workflows) |
| 🎵 **Music Generation** | `/v1/music/generations` (ComfyUI workflows) |
| 🛡️ **Moderations** | `/v1/moderations` safety checks |
| 🔀 **Reranking** | `/v1/rerank` for relevance scoring |
| 🔍 **Web Search** 🆕 | `/v1/search` — 5 providers (Serper, Brave, Perplexity, Exa, Tavily), 6,500+ free/month, auto-failover, cache |
### 🛡️ Resilience, Security & Governance
+374
View File
@@ -0,0 +1,374 @@
# ZWS_README_V4 — 启动性能优化:HMR 泄漏修复与 Turbopack 迁移
## 一、如何发现问题
### 现象
- `npm run dev` 后,首次打开浏览器白屏等待 **5-22 秒**不等。
- 运行一段时间后 Node 进程内存飙升至 **2.4 GB**,触发 Next.js 内存阈值保护强制重启。
- 重启后 `Ready in 82.6s`(正常冷启动仅 3.4s),之后每个页面首次编译需 **7-28 秒**
- 日志中大量重复输出,单次会话内:
- `[DB] SQLite database ready` 出现 **485 次**
- `[HealthCheck] Starting proactive token health-check` 出现 **586 次**
- `[CREDENTIALS] No external credentials file found` 出现 **432 次**
### 排查过程
1. **Terminal 日志分析**:统计关键日志出现次数,发现 DB 连接和 HealthCheck 定时器被反复创建。
2. **代码审计**:追踪到所有受影响模块使用 `let initialized = false` 作为单例守卫——这在 Next.js dev 模式的 Webpack HMR 下会被重置。
3. **对比**`apiBridgeServer.ts` 使用了 `globalThis.__omnirouteApiBridgeStarted`,在日志中无重复初始化,验证了 `globalThis` 方案的有效性。
4. **内存快照**:通过 `Get-Process node` 观察到两个 node 进程分别占用 1.7GB 和 1.0GB。
5. **编译时间分析**:日志中 `compile:` 字段显示 Webpack 编译每个路由需 2-26 秒,对比 Turbopack 应在 0.5-3 秒。
---
## 二、根因分析
### 根因 1(P0):模块级单例在 HMR 中丢失
Next.js dev 模式下,Webpack HMR 会重新执行被修改(或依赖链变化)的模块。模块级 `let` 变量在每次重新执行时被重置为初始值。
```typescript
// 修复前 — 每次 HMR 重新执行时 _db 重置为 null
let _db: SqliteDatabase | null = null;
export function getDbInstance() {
if (_db) return _db; // HMR 后这里永远 false
// ... 重新打开一个新的 DB 连接(旧连接泄漏)
}
```
**受影响的模块与泄漏类型:**
| 模块 | 泄漏资源 | 累计次数 | 后果 |
| ----------------------- | ---------------------- | -------- | ----------------------- |
| `db/core.ts` | SQLite 连接 | 485 | 文件句柄泄漏 + 内存占用 |
| `tokenHealthCheck.ts` | `setInterval` 定时器 | 586 | CPU 空转 + DB 查询风暴 |
| `localHealthCheck.ts` | `setTimeout` 定时器链 | ~400 | 重复 HTTP 请求 + CPU |
| `consoleInterceptor.ts` | console 方法包装 | ~400 | 日志 double-write |
| `gracefulShutdown.ts` | SIGTERM/SIGINT handler | ~400 | 信号处理器堆叠 |
**级联效应**:泄漏的资源持续消耗内存和 CPU → 触发 Next.js 内存阈值保护 → 进程重启 → Webpack 从零重建模块图 → **Ready in 82.6s**
### 根因 2P0):强制使用 Webpack 而非 Turbopack
`scripts/run-next.mjs` 中硬编码了 `--webpack` 标志:
```javascript
if (mode === "dev") {
args.splice(2, 0, "--webpack");
}
```
Next.js 16 默认使用 TurbopackRust 编写的增量打包器),dev 编译速度是 Webpack 的 5-10 倍。强制回退到 Webpack 导致:
| 指标 | Webpack | Turbopack(预期) |
| ----------------------- | ------- | ----------------- |
| 首页编译 | 3.7s | ~0.5s |
| Provider 详情页首次编译 | 22s | ~2-3s |
| API route 首次编译 | 2-7s | ~0.3-1s |
| 内存重启后 Ready | 82.6s | 不会触发 |
### 根因 3P1):`node:crypto` 被拉入客户端 bundle
`src/lib/db/proxies.ts` 使用了 `import { randomUUID } from "node:crypto"`。通过 `localDb.ts` 的 re-export 链,这个 Node.js 原生模块被间接拉入客户端组件的 bundle,导致 Webpack 报错:
```
UnhandledSchemeError: Reading from "node:crypto" is not handled by plugins
Import trace: node:crypto → ./src/lib/db/proxies.ts → ./src/lib/localDb.ts → page.tsx
```
Webpack 无法处理 `node:` URI scheme 前缀。`crypto`(不带 `node:` 前缀)已在 `next.config.mjs``serverExternalPackages` 中声明为服务端外部包。
### 根因 4P1):Edge Runtime 编译警告刷屏
Next.js 16 会同时为 **Node.js****Edge** 两种运行时编译 `instrumentation.ts`。虽然 `register()` 函数内有 `process.env.NEXT_RUNTIME === "nodejs"` 的运行时守卫,但 Turbopack 在打包 Edge 版本时仍会**静态追踪**所有动态 `import()` 的依赖链:
```
instrumentation.ts
→ import("@/lib/db/secrets")
@/lib/db/core.ts → fs, path, better-sqlite3
@/lib/dataPaths.ts → path, os
@/lib/db/migrationRunner.ts → fs, path, url
```
对每个 Node.js 原生模块,Turbopack 都输出一条 "not supported in Edge Runtime" 警告。每次有新请求触发热编译时,这组 **10+ 条警告重复刷一遍**,严重污染终端输出,干扰开发调试。
### 根因 5P2):启动 import 完全串行
`instrumentation.ts` 中 9 个 `await import()` 完全串行执行,每个都可能触发 Webpack 编译其依赖树:
```typescript
await ensureSecrets(); // 串行 1
const { initConsoleInterceptor } = await import(...); // 串行 2
const { initGracefulShutdown } = await import(...); // 串行 3
const { initApiBridgeServer } = await import(...); // 串行 4
const { startBackgroundRefresh } = await import(...); // 串行 5
const { getSettings } = await import(...); // 串行 6
const { setCustomAliases } = await import(...); // 串行 7
const { setDefaultFastServiceTierEnabled } = await import(...); // 串行 8
const { initAuditLog, cleanupExpiredLogs } = await import(...); // 串行 9
```
其中 4-6 互不依赖,7-8 互不依赖,完全可以并行。
---
## 三、修复方案
### 修复 1globalThis 单例守卫(core.ts, tokenHealthCheck.ts, localHealthCheck.ts, consoleInterceptor.ts, gracefulShutdown.ts
**原理**`globalThis` 对象在 Node.js 进程生命周期内全局唯一,不受 Webpack 模块重新执行的影响。
```typescript
// 修复后 — globalThis 在 HMR 后依然保留
declare global {
var __omnirouteDb: import("better-sqlite3").Database | undefined;
}
function getDb() {
return globalThis.__omnirouteDb ?? null;
}
function setDb(db) {
/* ... */
}
export function getDbInstance() {
const existing = getDb();
if (existing) return existing; // HMR 后命中缓存
// ...
}
```
**每个模块的具体改动:**
| 模块 | globalThis key | 守卫内容 |
| ----------------------- | ----------------------------------- | ----------------------------------------------------------- |
| `db/core.ts` | `__omnirouteDb` | SQLite 连接实例 |
| `tokenHealthCheck.ts` | `__omnirouteTokenHC` | `{ initialized, interval }` |
| `localHealthCheck.ts` | `__omnirouteLocalHC` | `{ initialized, sweepTimer, healthCache, sweepInProgress }` |
| `consoleInterceptor.ts` | `__omnirouteConsoleInterceptorInit` | `boolean` |
| `gracefulShutdown.ts` | `__omnirouteShutdownInit` | `boolean` |
**优点**
- 零依赖,无需额外库。
- 与 `apiBridgeServer.ts` 已有模式一致。
- 对生产环境零影响(非 HMR 场景下行为完全相同)。
**缺点/注意**
- `globalThis` 键名需全局唯一,使用 `__omniroute` 前缀避免冲突。
- 需要 `declare global` 类型声明以保持 TypeScript 类型安全。
- 生产构建中 `globalThis` 存储略冗余(但仅是一个对象引用,几乎零开销)。
### 修复 2:支持通过环境变量切换 Turbopackrun-next.mjs
```javascript
// 修复后 — 默认仍用 webpack(保持原有行为),设置环境变量可启用 Turbopack
if (mode === "dev" && process.env.OMNIROUTE_USE_TURBOPACK !== "1") {
args.splice(2, 0, "--webpack");
}
```
**默认行为不变**:dev 模式仍使用 Webpack,与修复前完全一致。设置 `OMNIROUTE_USE_TURBOPACK=1` 可切换到 Turbopack 以获得更快的 dev 编译速度。
**优点**
- 零风险:不改变任何人的现有体验。
- 需要时设置 `OMNIROUTE_USE_TURBOPACK=1` 即可获得 5-10 倍编译加速。
- `next.config.mjs` 中已有 `turbopack.resolveAlias` 配置,说明项目已在准备 Turbopack 迁移。
**缺点/注意**
- Turbopack 对某些 Webpack 特定配置(如自定义 externals 函数)的支持方式不同,启用前需测试兼容性。
- 默认走 Webpack 意味着不主动启用 Turbopack 的用户无法享受编译加速。
### 修复 3`node:crypto``crypto`proxies.ts, errorResponse.ts
```typescript
// 修复前
import { randomUUID } from "node:crypto";
// 修复后
import { randomUUID } from "crypto";
```
**优点**
- `crypto`(无 `node:` 前缀)已在 `next.config.mjs``serverExternalPackages` 列表中,Webpack/Turbopack 会正确将其标记为外部包。
- 消除 `UnhandledSchemeError` 构建失败。
- Node.js 中 `crypto``node:crypto` 解析到同一模块。
**缺点**
- 无。`crypto` 是 Node.js 内建模块,两种写法功能完全等价。
### 修复 4:分离 Edge/Node.js Instrumentationinstrumentation.ts → instrumentation-node.ts
**问题**`instrumentation.ts` 中所有 Node.js 逻辑(`ensureSecrets`、DB 初始化、审计日志等)虽然只在 `NEXT_RUNTIME === "nodejs"` 时执行,但 Turbopack 编译 Edge 版本时仍静态追踪其 import 链,对每个 `fs`/`path`/`os`/`better-sqlite3` 等原生模块输出警告。
**方案**:将所有 Node.js 专属逻辑提取到 `src/instrumentation-node.ts`,主文件通过**计算的 import 路径**引入,阻止 Turbopack 静态解析:
```typescript
// src/instrumentation.ts — 精简后仅 ~20 行
export async function register() {
if (process.env.NEXT_RUNTIME === "nodejs") {
// 拼接路径阻止 Turbopack 在 Edge 编译时静态解析模块依赖
const nodeMod = "./instrumentation-" + "node";
const { registerNodejs } = await import(nodeMod);
await registerNodejs();
}
}
```
```typescript
// src/instrumentation-node.ts — 包含全部 Node.js 启动逻辑
export async function registerNodejs(): Promise<void> {
await ensureSecrets();
// initConsoleInterceptor, initGracefulShutdown, initApiBridgeServer, ...
// (原 instrumentation.ts 的完整 Node.js 逻辑)
}
```
**关键技术**`"./instrumentation-" + "node"` 是运行时拼接的字符串,Turbopack 无法在编译期确定其值,因此**不会追踪**该 import 的依赖树。Node.js 运行时则正常解析该路径并执行。
**优点**
- Edge 编译时完全跳过 Node.js 模块追踪,**10+ 条重复警告全部消除**。
- Node.js 运行时行为与修复前完全一致。
- 启动时间从 **13.9s → 1.25s**Turbopack 不再在 Edge 编译中处理 Node.js 模块图)。
**缺点/注意**
- 新增一个文件 `instrumentation-node.ts`,需同步维护。
- 计算 import 路径是有意为之的 bundler 逃逸技巧,需加注释说明原因防止后续重构时被"优化"回静态字符串。
### 修复 5:并行化 instrumentation.ts 中的启动 import
```typescript
// 修复后 — 4 个独立模块并行导入
const [
{ initGracefulShutdown },
{ initApiBridgeServer },
{ startBackgroundRefresh },
{ getSettings },
] = await Promise.all([
import("@/lib/gracefulShutdown"),
import("@/lib/apiBridgeServer"),
import("@/domain/quotaCache"),
import("@/lib/db/settings"),
]);
// 2 个 open-sse 模块也并行导入
const [{ setCustomAliases }, { setDefaultFastServiceTierEnabled }] = await Promise.all([
import("@omniroute/open-sse/services/modelDeprecation.ts"),
import("@omniroute/open-sse/executors/codex.ts"),
]);
```
**优点**
- `consoleInterceptor` 仍保持第一个(必须在任何日志前初始化)。
- 后续 4 个无依赖模块并行加载,节省 3 次串行等待。
- open-sse 的 2 个模块也并行加载。
**缺点**
- 并行 import 的错误堆栈略复杂(Promise.all 中某一个失败会 reject 整个组)。
- 这里的 compliance 模块仍保持独立 try/catch 串行,因为它有自己的错误处理逻辑。
---
## 四、预期效果
| 指标 | 修复前 | 修复后(预期) |
| ----------------------------- | ------------------------- | ------------------------ |
| DB 连接创建次数 | 485 次/会话 | 1 次 |
| HealthCheck 定时器 | 586 个泄漏 | 1 个 |
| 信号处理器注册 | ~400 次重复 | 1 次 |
| Console 拦截层数 | ~400 层嵌套 | 1 层 |
| 内存使用峰值 | 2.4 GB → OOM 重启 | 预期 < 500 MB |
| 冷启动 Ready | 3.4s | ~3s(略快) |
| 内存重启 Ready | 82.6s | 不再触发内存重启 |
| Login 页首次编译 | 3.7s | ~0.5s (需启用 Turbopack) |
| Provider 详情页首次编译 | 22s | ~2-3s (需启用 Turbopack) |
| `node:crypto` 构建错误 | 反复出现 | 消除 |
| Edge Runtime 编译警告 | 每次热编译刷出 10+ 条 | **0 条** |
| instrumentation 启动耗时 | 13.9s(含 Edge 模块追踪) | **1.25s** |
| instrumentation import 并行度 | 9 次串行 import | 3 批并行 import |
---
## 五、涉及文件清单
| 区域 | 文件 | 改动类型 |
| ------------------- | ------------------------------- | ------------------------------------------------------------------ |
| DB 单例 | `src/lib/db/core.ts` | `let _db``globalThis.__omnirouteDb` |
| Token 健康检查 | `src/lib/tokenHealthCheck.ts` | `let initialized``globalThis.__omnirouteTokenHC` |
| 本地节点健康检查 | `src/lib/localHealthCheck.ts` | `let initialized``globalThis.__omnirouteLocalHC` |
| Console 拦截 | `src/lib/consoleInterceptor.ts` | `let initialized``globalThis.__omnirouteConsoleInterceptorInit` |
| 优雅关停 | `src/lib/gracefulShutdown.ts` | 新增 `globalThis.__omnirouteShutdownInit` 守卫 |
| Dev 启动脚本 | `scripts/run-next.mjs` | 新增 `OMNIROUTE_USE_TURBOPACK=1` 开关 |
| Proxy 注册表 | `src/lib/db/proxies.ts` | `node:crypto``crypto` |
| API 错误响应 | `src/lib/api/errorResponse.ts` | `node:crypto``crypto` |
| 启动钩子(主入口) | `src/instrumentation.ts` | 精简为 ~20 行,计算 import 路径阻止 Edge 追踪 |
| 启动钩子(Node.js | `src/instrumentation-node.ts` | 新文件,承载全部 Node.js 启动逻辑 + `Promise.all` 并行 |
---
## 六、回退方案
- **启用 Turbopack**:设置 `OMNIROUTE_USE_TURBOPACK=1` 环境变量;不设置则默认使用 Webpack(原有行为不变)。
- **globalThis 方案异常**:所有 globalThis key 都以 `__omniroute` 为前缀,可通过 `delete globalThis.__omnirouteDb` 等方式手动重置。
- **Edge 警告回退**:若 `instrumentation-node.ts` 拆分导致问题,可将其内容合并回 `instrumentation.ts`,恢复为直接 `import()` 调用(警告会重新出现但不影响功能)。
- **生产环境**:以上修复对生产构建无负面影响——生产环境不存在 HMR,globalThis 单例仅在首次调用时初始化一次。计算 import 路径在 `next build` 时由 Node.js 正常解析,不影响打包产物。
---
## 七、单元测试与备份恢复(pre-commit 验证通过)
为保证提交前必须通过验证(不再使用 `--no-verify`),对以下失败用例与生产逻辑做了修复与加固。
### 问题与根因
| 失败项 | 根因 |
| ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
| bootstrap-env 4 个用例 | Windows 上 DATA_DIR 解析用 `APPDATA`/`homedir()`,测试只设了 `HOME`,脚本读不到测试用的 `.env`。 |
| domain-persistence costRules 2 个用例 | `core` 在首次 import 时缓存 `DATA_DIR`;测试每测一个 tmpDir 并在 afterEach 删目录,导致后续 describe 使用的 DB 路径已被删,读写得到 0。 |
| fixes-p1 restoreDbBackup | 测试在 DB 仍打开时写 stale 侧文件;`restoreDbBackup` 内 pre-restore 备份未 await 就关库,Windows 上句柄未及时释放,unlink 报 EBUSY。 |
| fixes-p1 resetStorage 及后续用例 | 上一测留下 DB 打开,下一测 `resetStorage()` 删目录时文件仍被占用,EBUSY。 |
### 修复 6bootstrap-env 测试(tests/unit/bootstrap-env.test.mjs
在每个用例的 `withTempEnv` 回调开头增加 `process.env.DATA_DIR = dataDir`,使脚本在任意平台(含 Windows)都使用测试临时目录,而不是依赖 `HOME`/`APPDATA`
### 修复 7domain-persistence 测试(tests/unit/domain-persistence.test.mjs
- **单例 tmpDir**:全文件共用一个 `fileTmpDir`,在模块加载时创建并设置 `process.env.DATA_DIR`,与 `core` 首次加载时缓存的路径一致。
- **每测清 DB 不清目录**`beforeEach``resetDbInstance()` 后删除 `storage.sqlite` 及其 `-wal`/`-shm`/`-journal`,保证每测干净 DB,不在 afterEach 删目录,避免路径失效。
- **收尾**`after()` 中恢复 `DATA_DIR` 并删除 `fileTmpDir`
- **costRules 断言**:改为小容差精确校验(`assertAlmostEqual`),继续验证 `4.5` / `4.0` 这类业务关键值,避免把真实累计错误放过去。
### 修复 8fixes-p1 测试(tests/unit/fixes-p1.test.mjs
- **restoreDbBackup 用例**:在写入 stale 侧文件前调用 `core.resetDbInstance()`,避免 DB 仍打开时写 `-wal`/`-shm` 触发 Windows 锁错误。
- **Windows 跳过**:该用例在 Windows 上仍使用 `test(..., { skip: isWindows })`。原因不是业务逻辑不支持 Windows,而是 better-sqlite3 关闭后底层句柄释放存在时序抖动,这条真实 sidecar 集成测试容易退化成不稳定的文件锁测试;Linux/macOS 上照常运行。
- **核心兜底测试**:新增平台无关的 `unlinkFileWithRetry` 单测,直接模拟 `EBUSY` / `EPERM` 后重试并最终成功,确保 Windows 相关的重试删除逻辑被稳定覆盖,而不是完全依赖 flaky 的真实文件锁时序。
- **resetStorage**:改为 async,对 `rmSync(TEST_DATA_DIR)` 做最多 10 次、间隔 100ms 的 EBUSY/EPERM 重试,避免下一测因上一测句柄未释放而失败。
### 修复 9:备份恢复逻辑(src/lib/db/backup.ts
- **pre-restore 备份改为同步等待**:在 `restoreDbBackup` 内用内联逻辑做 pre-restore 备份并 `await` 完成,再调用 `resetDbInstance()`,避免异步 backup 未结束就关库导致后续 unlink 失败。
- **节流语义保持一致**pre-restore 备份成功后补回 `_lastBackupAt = Date.now()`,避免恢复后紧接着又触发一轮额外自动备份。
- **关库后短延迟**`resetDbInstance()``await new Promise(r => setTimeout(r, 500))`,再执行 unlink,给 Windows 等平台释放句柄留时间。
- **unlink 重试**:将主库及 `-wal`/`-shm`/`-journal` 的删除提取为 `unlinkFileWithRetry`,统一做最多 10 次、间隔 100ms 的 EBUSY/EPERM 重试,提高恢复流程在锁释放较慢环境下的成功率,也便于单测直接覆盖重试逻辑。
### 涉及文件(本节)
| 区域 | 文件 | 改动类型 |
| -------- | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
| 单元测试 | `tests/unit/bootstrap-env.test.mjs` | 各用例内设置 `process.env.DATA_DIR = dataDir` |
| 单元测试 | `tests/unit/domain-persistence.test.mjs` | 单例 tmpDir、beforeEach 清 DB 文件、after 删目录;costRules 改为小容差精确断言 |
| 单元测试 | `tests/unit/fixes-p1.test.mjs` | restoreDbBackup 前 resetDbInstance、Windows skip 说明、resetStorage 重试、`unlinkFileWithRetry` 核心单测 |
| 备份恢复 | `src/lib/db/backup.ts` | pre-restore 内联并 await、恢复 `_lastBackupAt` 节流语义、关库后 500ms 延迟、抽取 `unlinkFileWithRetry` 重试删除 |
+332
View File
@@ -0,0 +1,332 @@
# ZWS_README_V5 — 按协议配置模型兼容性 + 前端性能优化
V4 内容(HMR 泄漏修复、Edge 警告消除、测试稳定性)已完成;V5 在 V4 基础上实现**按协议维度配置模型兼容性**,新增前端查找性能优化与类型安全改进。
---
## 一、如何发现问题
### 现象
- 同一模型被 **OpenAI Chat Completions**、**OpenAI Responses API**、**Anthropic Messages** 三种客户端请求形态调用时,V2 的兼容性开关(工具 ID 9 位、不保留 developer 角色)是**全局生效**的——无法为不同协议设置不同的兼容策略。
- 例如:用户希望 OpenAI Responses API 请求时不保留 developer 角色(MiniMax 422 修复),但 OpenAI Chat Completions 请求时保留。V2 下只能二选一。
- 前端兼容性弹层未标明当前配置对应哪种协议,容易误导。
- 前端组件中 `Array.find()` 在每次渲染时对 customModels 和 modelCompatOverrides 做 O(n) 线性扫描,模型数量多时存在不必要的性能开销。
- `ModelCompatPatch` 类型定义与运行时逻辑不一致:`preserveOpenAIDeveloperRole` 字段需要支持 `null`(表示取消设置/恢复默认),但类型仅允许 `boolean`
### 排查过程
1. **需求分析**:梳理 `detectFormat(body)` 返回的三种协议键(`openai``openai-responses``claude`),确认每种协议对 developer 角色和 tool call ID 的需求不同。
2. **数据模型设计**:在现有 `normalizeToolCallId` / `preserveOpenAIDeveloperRole` 顶层字段基础上,设计 `compatByProtocol` 嵌套结构,按协议键细分。
3. **构建问题**:客户端 `"use client"` 组件直接从 `@/lib/localDb` 引入常量时,间接拉入了 `node:crypto`(经由 `db/proxies.ts`),触发 Webpack `UnhandledSchemeError`。需将常量拆到 `shared/` 层。
4. **前端性能**:通过 React DevTools 和代码审计发现 `effectiveNormalizeForProtocol` 等函数每次调用都对数组做 `find()`,在渲染列表时存在 O(n²) 的隐患。
---
## 二、根因分析
### 根因 1(P0):兼容选项无协议维度
V2 的 `normalizeToolCallId` / `preserveOpenAIDeveloperRole` 存储在模型级别的顶层字段,无法区分请求来源协议。`chatCore.ts` 中的 getter 函数只接收 `(providerId, modelId)` 两个参数,不感知当前请求的 `sourceFormat`
**影响**:跨协议场景下用户只能设置一个全局值,无法精确控制。
### 根因 2P1):客户端构建拉入 Node.js 模块
`page.tsx`"use client")→ `@/lib/localDb``db/proxies.ts``import { randomUUID } from "node:crypto"`
Webpack 无法处理 `node:` URI scheme,报 `UnhandledSchemeError`。虽然 V4 已将 `node:crypto``crypto` 修复了 `proxies.ts`,但 `localDb.ts` 的 barrel export 链仍然存在风险——客户端组件不应引入任何可能传递到 Node.js 模块的路径。
### 根因 3P2):前端查找性能
`effectiveNormalizeForProtocol``effectivePreserveForProtocol``anyNormalizeCompatBadge``anyNoPreserveCompatBadge` 四个函数每次调用都使用 `Array.find()``customModels``modelCompatOverrides` 数组中查找目标模型。在模型列表渲染时,每个模型行会调用多次这些函数,导致 O(n × m) 的查找开销(n = 模型数,m = 每行调用次数)。
### 根因 4(P2):类型定义与运行时不一致
```typescript
// V3 暂存区版本(有问题)
export type ModelCompatPatch = Partial<
Pick<
ModelCompatOverride,
"normalizeToolCallId" | "preserveOpenAIDeveloperRole" | "compatByProtocol"
>
>;
```
`ModelCompatOverride.preserveOpenAIDeveloperRole` 类型为 `boolean | undefined`,但 `mergeModelCompatOverride()` 内部有 `=== null` 判断(用于取消设置/恢复默认),类型层面无法覆盖。
---
## 三、修复方案
### 修复 1`compatByProtocol` 存储与读取(models.ts
**新增数据结构**
```typescript
type CompatByProtocolMap = Partial<Record<ModelCompatProtocolKey, ModelCompatPerProtocol>>;
export type ModelCompatOverride = {
id: string;
normalizeToolCallId?: boolean;
preserveOpenAIDeveloperRole?: boolean;
compatByProtocol?: CompatByProtocolMap; // 新增
};
```
**读取优先级链**(适用于 `getModelNormalizeToolCallId``getModelPreserveOpenAIDeveloperRole`):
```
compatByProtocol[sourceFormat].field → 顶层 field → 默认值
```
1. 若 `sourceFormat` 属于已知协议键(`openai` / `openai-responses` / `claude`),且 `compatByProtocol[sourceFormat]` 中存在目标字段,使用该值。
2. 否则回退到顶层字段。
3. 顶层字段也不存在时使用默认值(normalizeToolCallId=falsepreserveOpenAIDeveloperRole=undefined)。
**深度合并逻辑** `deepMergeCompatByProtocol()`
- 对每个协议键,逐字段合并而非覆盖。
- `normalizeToolCallId=false` 时删除该字段(不存储 false,减少冗余)。
- 合并后若整个协议条目为空对象,删除该协议条目。
- 协议键通过 `isCompatProtocolKey()` 白名单校验,拒绝未知键。
**Getter 签名扩展**(向后兼容,第三参数可选):
```typescript
export function getModelNormalizeToolCallId(
providerId: string,
modelId: string,
sourceFormat?: string | null
): boolean;
export function getModelPreserveOpenAIDeveloperRole(
providerId: string,
modelId: string,
sourceFormat?: string | null
): boolean | undefined;
```
**优点**
- 完全向后兼容:无 `sourceFormat` 参数时行为与 V2 完全一致。
- 协议键白名单校验防止存储污染。
- 深度合并保留未变更协议的配置。
**缺点/注意**
- JSON 存储体积略增(每个模型最多增加 3 个协议条目)。
- 新增 ~80 行 TypeScript 代码。
### 修复 2:请求管线传入 sourceFormatchatCore.ts
```typescript
const normalizeToolCallId = getModelNormalizeToolCallId(
provider || "",
model || "",
sourceFormat // 新增第三参
);
const preserveDeveloperRole = getModelPreserveOpenAIDeveloperRole(
provider || "",
model || "",
sourceFormat // 新增第三参
);
```
`sourceFormat` 由已有的 `detectFormat(body)` 返回,无需新增检测逻辑。
**优点**
- 改动仅 2 行,精准传参。
- 不影响其他 handlerembeddings、imageGeneration 等不涉及 developer 角色和 tool call ID)。
### 修复 3API 路由支持 compatByProtocolroute.ts
**PUT 请求体扩展**
- 解构 `compatByProtocol` 并传入 `updateCustomModel()`
- `compatOnly` 判断扩展:仅含 `provider` + `modelId` + 兼容字段时,走 `mergeModelCompatOverride()` 路径。
- 使用 `ModelCompatPatch` 类型替代行内类型定义,统一类型来源。
**Zod 校验 schema**
```typescript
const modelCompatPerProtocolSchema = z.object({
normalizeToolCallId: z.boolean().optional(),
preserveOpenAIDeveloperRole: z.boolean().optional(),
}).strict(); // strict: 拒绝额外字段
compatByProtocol: z
.record(z.enum(["openai", "openai-responses", "claude"]), modelCompatPerProtocolSchema)
.optional(),
```
**优点**
- `.strict()` 防止客户端注入额外字段。
- `z.enum()` 限定协议键,与后端白名单一致。
- 仅传 `compatByProtocol` 即可更新,前端无需拼装完整模型对象。
### 修复 4:客户端安全常量拆分(modelCompat.ts
**新增** `src/shared/constants/modelCompat.ts`
```typescript
export const MODEL_COMPAT_PROTOCOL_KEYS = ["openai", "openai-responses", "claude"] as const;
export type ModelCompatProtocolKey = (typeof MODEL_COMPAT_PROTOCOL_KEYS)[number];
```
- 不依赖 Node.js / DB 代码,客户端组件可安全引入。
- `models.ts` 从此模块引入并再导出。
- `localDb.ts` 新增 `ModelCompatPatch` 类型导出(供 route.ts 使用),不导出协议常量。
- `page.tsx` 改为从 `@/shared/constants/modelCompat` 引入。
**优点**
- 彻底切断客户端 → localDb → db → proxies → node:crypto 的依赖链。
- 协议键定义单一来源(Single Source of Truth)。
### 修复 5:前端协议选择器与按协议解析(page.tsx)
**ModelCompatPopover 重构**
- 新增协议下拉选择器(`<select>`),可选 OpenAI Chat / OpenAI Responses / Anthropic Messages。
- 两个开关(工具 ID 9 位、不保留 developer**针对选中协议**生效。
- 选择 Claude 协议时隐藏 developer 角色开关(developer 仅对 OpenAI 系有意义)。
- 保存时以 `{ compatByProtocol: { [protocol]: payload } }` 形式提交,后端按协议合并。
- 深色模式适配:下拉框使用 `bg-white dark:bg-zinc-800``text-zinc-900 dark:text-zinc-100`
**Props 接口重构**
旧接口(4 个独立值/回调):
```typescript
(normalizeToolCallId, preserveDeveloperRole, onNormalizeChange, onPreserveChange);
```
新接口(3 个函数式 props):
```typescript
effectiveModelNormalize: (protocol: string) => boolean
effectiveModelPreserveDeveloper: (protocol: string) => boolean
onCompatPatch: (protocol: string, payload: {...}) => void
```
所有消费方(`ModelRow``PassthroughModelRow``CustomModelsSection``CompatibleModelsSection`)已同步更新。
**角标显示逻辑**
- `anyNormalizeCompatBadge()`:任意协议或顶层存在 `normalizeToolCallId=true` 即显示「ID×9」角标。
- `anyNoPreserveCompatBadge()`:任意协议或顶层存在 `preserveOpenAIDeveloperRole=false` 即显示「不保留」角标。
**CustomModelsSection 增强**
- 新增 `modelCompatOverrides` 状态,从 API 响应中获取。
- 新增 `saveCustomCompat()` 函数,支持仅传 `compatByProtocol` 的独立保存。
### 修复 6:前端 Map 查找性能优化(page.tsx
**问题**`effectiveNormalizeForProtocol` 等函数对 `customModels``modelCompatOverrides``Array.find()` 做 O(n) 查找,在列表渲染时每个模型行多次调用。
**方案**:使用 `useMemo` + `Map` 将数组预建为 O(1) 查找表。
```typescript
type CompatModelMap = Map<string, CompatModelRow>;
function buildCompatMap(rows: CompatModelRow[]): CompatModelMap {
const m = new Map<string, CompatModelRow>();
for (const r of rows) if (r.id) m.set(r.id, r);
return m;
}
// 在组件内
const customMap = useMemo(() => buildCompatMap(modelMeta.customModels), [modelMeta.customModels]);
const overrideMap = useMemo(
() => buildCompatMap(modelMeta.modelCompatOverrides),
[modelMeta.modelCompatOverrides]
);
```
所有查找函数签名从 `(modelId, protocol, customModels[], overrides[])` 改为 `(modelId, protocol, customMap, overrideMap)`,内部使用 `Map.get()` 替代 `Array.find()`
**优点**
- 查找从 O(n) 降为 O(1)。
- `useMemo` 依赖项正确,仅在数据变化时重建 Map。
- `CustomModelsSection` 内部也独立构建 Map,不依赖父组件。
### 修复 7ModelCompatPatch 类型修正(models.ts
```typescript
// 修复后 — 显式允许 null
export type ModelCompatPatch = {
normalizeToolCallId?: boolean;
preserveOpenAIDeveloperRole?: boolean | null; // null = 取消设置/恢复默认
compatByProtocol?: CompatByProtocolMap;
};
```
`mergeModelCompatOverride()` 内的 `=== null` 判断逻辑一致,类型安全。
### 修复 8CompatByProtocolMap 类型收紧(page.tsx
客户端 `CompatByProtocolMap``Record<string, ...>` 改为 `Record<ModelCompatProtocolKey, ...>`,增强类型安全,防止传入未知协议键。
### 修复 9i18n 文案新增
| 键名 | 中文 | 英文 |
| ------------------------------- | --------------------------------------------- | -------------------------------------------------------------- |
| `compatProtocolLabel` | 客户端请求协议 | Client request protocol |
| `compatProtocolHint` | 以下选项在 OmniRoute 识别到该请求形态时生效。 | These options apply when OmniRoute detects this request shape. |
| `compatProtocolOpenAI` | OpenAI Chat Completions | OpenAI Chat Completions |
| `compatProtocolOpenAIResponses` | OpenAI Responses API | OpenAI Responses API |
| `compatProtocolClaude` | Anthropic Messages | Anthropic Messages |
---
## 四、使用方式
1. 点击模型行的 **「兼容性」** 按钮。
2. 在弹层内先选择 **「客户端请求协议」**OpenAI Chat / OpenAI Responses / Anthropic Messages)。
3. 勾选该协议下的「工具 ID 9 位」或「不保留 developer 角色」。
4. 保存后,仅在该协议形态的请求下生效。
5. 未配置某协议时,该协议下行为回退到顶层兼容字段(若存在),再回退到默认值(保留 developer、不规范化 tool id)。
6. 角标「ID×9」「不保留」在任意协议存在对应配置时显示。
---
## 五、预期效果
| 指标 | 修复前 | 修复后 |
| ------------------------- | --------------------- | ------------------------------------------ |
| 兼容性配置维度 | 全局(模型级) | 按协议(OpenAI Chat / Responses / Claude |
| developer 角色精确控制 | 不支持 | 支持(如:仅 Responses API 不保留) |
| 前端兼容性查找性能 | O(n) Array.find | O(1) Map.getuseMemo 缓存) |
| ModelCompatPatch 类型安全 | null 值无类型覆盖 | 显式 `boolean \| null` |
| 客户端构建风险 | 可能引入 Node.js 模块 | 已隔离(shared/constants 层) |
| API 验证 | 无 compatByProtocol | Zod strict schema 校验 |
| 深色模式 | 协议选择器不可读 | bg/text 适配 dark 主题 |
---
## 六、涉及文件清单
| 区域 | 文件 | 改动类型 |
| ---------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
| 协议常量 | `src/shared/constants/modelCompat.ts` | **新建**,客户端安全的协议键与类型 |
| 存储与读写 | `src/lib/db/models.ts` | `compatByProtocol` 数据结构、深度合并、getter 第三参 `sourceFormat``ModelCompatPatch` 类型修正 |
| 再导出层 | `src/lib/localDb.ts` | 新增 `ModelCompatPatch` 类型导出 |
| API 路由 | `src/app/api/provider-models/route.ts` | PUT 支持 `compatByProtocol`,使用 `ModelCompatPatch` 类型 |
| 输入校验 | `src/shared/validation/schemas.ts` | `modelCompatPerProtocolSchema`strict+ `compatByProtocol` 记录校验 |
| 请求管线 | `open-sse/handlers/chatCore.ts` | `getModelNormalizeToolCallId` / `getModelPreserveOpenAIDeveloperRole` 传入 `sourceFormat` |
| 前端 UI | `src/app/(dashboard)/dashboard/providers/[id]/page.tsx` | 协议选择器、按协议解析/保存、角标逻辑、Map 性能优化、类型收紧 |
| i18n | `src/i18n/messages/en.json``src/i18n/messages/zh-CN.json` | 5 条新文案 |
---
## 七、回退方案
- **禁用按协议配置**:删除 `compatByProtocol` 字段后,getter 自动回退到顶层字段,行为与 V2 一致。
- **前端 Map 优化回退**:将 `Map.get()` 改回 `Array.find()` 即可,纯性能优化无功能耦合。
- **客户端常量回退**:将 `MODEL_COMPAT_PROTOCOL_KEYS` 定义移回 `models.ts` 并从 `localDb.ts` 导出(需同时确保 `node:crypto` 问题不再存在)。
- **生产环境**:以上修复对生产构建无负面影响。`compatByProtocol` 为可选字段,未配置时默认行为不变。API Zod 校验确保不会接受畸形数据。
+21 -2
View File
@@ -189,8 +189,27 @@ const serverJs = join(APP_DIR, "server.js");
if (!existsSync(serverJs)) {
console.error("\x1b[31m✖ Server not found at:\x1b[0m", serverJs);
console.error(" This usually means the package was not built correctly.");
console.error(" Try reinstalling: npm install -g omniroute");
console.error(" The package may not have been built correctly.");
console.error("");
// (#492) Detect common non-standard Node managers that cause this issue
const nodeExec = process.execPath || "";
const isMise = nodeExec.includes("mise") || nodeExec.includes(".local/share/mise");
const isNvm = nodeExec.includes(".nvm") || nodeExec.includes("nvm");
if (isMise) {
console.error(
" \x1b[33m⚠ mise detected:\x1b[0m If you installed via `npm install -g omniroute`,"
);
console.error(" try: \x1b[36mnpx omniroute@latest\x1b[0m (downloads a fresh copy)");
console.error(" or: \x1b[36mmise exec -- npx omniroute\x1b[0m");
} else if (isNvm) {
console.error(
" \x1b[33m⚠ nvm detected:\x1b[0m Try reinstalling after loading the correct Node version:"
);
console.error(" \x1b[36mnvm use --lts && npm install -g omniroute\x1b[0m");
} else {
console.error(" Try: \x1b[36mnpm install -g omniroute\x1b[0m (reinstall)");
console.error(" Or: \x1b[36mnpx omniroute@latest\x1b[0m");
}
process.exit(1);
}
+2 -2
View File
@@ -932,8 +932,8 @@ npm run electron:build:linux # Linux (.AppImage)
| ميزة | ماذا يفعل || -------------------------- | ------------------------------------------------------------- |
| 🖼️ **إنشاء الصور** | `/v1/images/generations` مع الواجهات الخلفية السحابية والمحلية |
| 📐 **المضامين** | `/v1/embeddings` للبحث وخطوط أنابيب RAG |
| 🎤 **نسخ صوتي** | `/v1/audio/transcriptions` (مقدمو خدمات الهمس والإضافيون) |
| 🔊 **تحويل النص إلى كلام** | `/v1/audio/speech` (محركات/موفرو متعددون) |
| 🎤 **نسخ صوتي** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **تحويل النص إلى كلام** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🎬 **توليد الفيديو** | `/v1/videos/generations` (سير عمل ComfyUI + SD WebUI) |
| 🎵 **جيل الموسيقى** | `/v1/music/generations` (سير عمل ComfyUI) |
| 🛡️ **اعتدالات** | فحوصات السلامة `/v1/moderations` |
+2 -2
View File
@@ -933,8 +933,8 @@ OmniRoute v2.0 е създаден като операционна платфо
| Характеристика | Какво прави || -------------------------- | ------------------------------------------------------------ |
| 🖼️ **Генериране на изображения** | `/v1/images/generations` с облак и локален бекенд |
| 📐 **Вграждания** | `/v1/embeddings` за търсене и RAG тръбопроводи |
| 🎤 **Аудио транскрипция** | `/v1/audio/transcriptions` (Whisper и допълнителни доставчици) |
| 🔊 **Текст към говор** | `/v1/audio/speech` (множество машини/доставчици) |
| 🎤 **Аудио транскрипция** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Текст към говор** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🎬 **Видео генериране** | `/v1/videos/generations` (работни процеси ComfyUI + SD WebUI) |
| 🎵 **Музикално поколение** | `/v1/music/generations` (работни процеси на ComfyUI) |
| 🛡️ **Модерации** | `/v1/moderations` проверки за безопасност |
+225 -349
View File
@@ -2,7 +2,7 @@
### Nikdy nepřestávejte s kódováním. Chytré směrování k **BEZPLATNÝM a levným modelům AI** s automatickým přepínáním mezi záložními systémy.
*Váš univerzální API proxy jeden endpoint, více než 44 poskytovatelů, nulové výpadky. Nyní s orchestrací agentů **MCP a A2A** .*
_Váš univerzální API proxy jeden endpoint, více než 44 poskytovatelů, nulové výpadky. Nyní s orchestrací agentů **MCP a A2A** ._
**Dokončení chatu • Vkládání • Generování obrázků • Video • Hudba • Audio • Změna pořadí • **Vyhledávání na webu** • MCP server • A2A protokol • 100% TypeScript**
@@ -30,26 +30,23 @@
<summary><b>Kliknutím zobrazíte snímky obrazovky z řídicího panelu</b></summary>
</details>
Strana | Snímek obrazovky
--- | ---
**Poskytovatelé** | ![Poskytovatelé](docs/screenshots/01-providers.png)
**Kombinace** | ![Kombinace](docs/screenshots/02-combos.png)
**Analytika** | ![Analytika](docs/screenshots/03-analytics.png)
**Zdraví** | ![Zdraví](docs/screenshots/04-health.png)
**Překladatel** | ![Překladatel](docs/screenshots/05-translator.png)
**Nastavení** | ![Nastavení](docs/screenshots/06-settings.png)
**Nástroje CLI** | ![Nástroje CLI](docs/screenshots/07-cli-tools.png)
**Protokoly používání** | ![Používání](docs/screenshots/08-usage.png)
**Koncové body** | ![Koncové body](docs/screenshots/09-endpoint.png)
| Strana | Snímek obrazovky |
| ----------------------- | --------------------------------------------------- |
| **Poskytovatelé** | ![Poskytovatelé](docs/screenshots/01-providers.png) |
| **Kombinace** | ![Kombinace](docs/screenshots/02-combos.png) |
| **Analytika** | ![Analytika](docs/screenshots/03-analytics.png) |
| **Zdraví** | ![Zdraví](docs/screenshots/04-health.png) |
| **Překladatel** | ![Překladatel](docs/screenshots/05-translator.png) |
| **Nastavení** | ![Nastavení](docs/screenshots/06-settings.png) |
| **Nástroje CLI** | ![Nástroje CLI](docs/screenshots/07-cli-tools.png) |
| **Protokoly používání** | ![Používání](docs/screenshots/08-usage.png) |
| **Koncové body** | ![Koncové body](docs/screenshots/09-endpoint.png) |
---
### 🤖 Bezplatný poskytovatel umělé inteligence pro vaše oblíbené programátory
*Připojte libovolný nástroj IDE nebo CLI s umělou inteligencí přes OmniRoute — bezplatnou API bránu pro neomezené kódování.*
_Připojte libovolný nástroj IDE nebo CLI s umělou inteligencí přes OmniRoute — bezplatnou API bránu pro neomezené kódování._
<table>
<tr>
@@ -68,7 +65,6 @@ Strana | Snímek obrazovky
</tr>
</table>
<sub>📡 Všichni agenti se připojují přes <code>http://localhost:20128/v1</code> nebo <code>http://cloud.omniroute.online/v1</code> — jedna konfigurace, neomezené modely a kvóty</sub>
---
@@ -161,9 +157,6 @@ Vývojáři platí za Claude Pro, Codex Pro nebo GitHub Copilot 20200 dolarů
- **Vlastní kombinace** — Přizpůsobitelné záložní řetězce se 6 strategiemi vyvažování (fill-first, round robin, P2C, náhodné, nejméně používané, nákladově optimalizované)
- **Codex Business Quotas** — Sledování kvót pracovního prostoru firmy/týmu přímo v dashboardu
<details>
<summary><b>🔌 2. „Potřebuji použít více poskytovatelů, ale každý má jiné API“</b></summary>
</details>
@@ -180,9 +173,6 @@ OpenAI používá jeden formát, Claude (Anthropic) jiný a Gemini ještě třet
- **Strukturovaný výstup pro Gemini**`json_schema` → automatická konverze `responseMimeType` / `responseSchema`
- **Výchozí hodnota `stream` je `false`** Odpovídá specifikaci OpenAI, čímž se zabrání neočekávanému SSE v Python/Rust/Go SDK.
<details>
<summary><b>🌐 3. „Můj poskytovatel AI blokuje můj region/zemi“</b></summary>
</details>
@@ -199,9 +189,6 @@ Poskytovatelé jako OpenAI/Codex blokují přístup z určitých geografických
- **TLS Fingerprint Spoofing** — Otisk prstu TLS podobný prohlížeči pomocí `wreq-js` pro obcházení detekce botů
- **🔏 Porovnávání otisků prstů v CLI** — Změní pořadí záhlaví a polí v těle serveru tak, aby odpovídala nativním binárním podpisům v CLI, čímž drasticky snižuje riziko nahlašování účtu. IP adresa proxy je zachována — získáte současně stealth **i** maskování IP adresy.
<details>
<summary><b>🆓 4. „Chci používat umělou inteligenci pro kódování, ale nemám peníze“</b></summary>
</details>
@@ -216,9 +203,6 @@ Ne každý si může dovolit zaplatit 20200 dolarů měsíčně za předplatn
- **NVIDIA NIM Free Access** — ~40 RPM developerský přístup k více než 70 modelům na build.nvidia.com (přechod z kreditů na čisté limity rychlosti)
- **Strategie optimalizace nákladů** Strategie směrování, která automaticky vybere nejlevnějšího dostupného poskytovatele
<details>
<summary><b>🔒 5. „Potřebuji chránit svou bránu umělé inteligence před neoprávněným přístupem“</b></summary>
</details>
@@ -236,9 +220,6 @@ Při zpřístupnění brány umělé inteligence síti (LAN, VPS, Docker) může
- **Ochrana proti vkládání výzev** Sanitizace proti škodlivým vzorcům výzev
- **Šifrování AES-256-GCM** přihlašovací údaje jsou v klidovém stavu šifrovány
<details>
<summary><b>🛑 6. „Můj poskytovatel selhal a já ztratil/a programovací tok“</b></summary>
</details>
@@ -254,9 +235,6 @@ Poskytovatelé umělé inteligence se mohou stát nestabilními, vracet chyby 5x
- **Kombinovaný jistič** Automaticky deaktivuje selhávajícího poskytovatele v rámci kombinovaného řetězce
- **Dashboard stavu** — Monitorování provozuschopnosti, stavy jističů, uzamčení, statistiky mezipaměti, latence p50/p95/p99
<details>
<summary><b>🔧 7. „Konfigurace každého nástroje umělé inteligence je zdlouhavá a opakující se“</b></summary>
</details>
@@ -270,9 +248,6 @@ Vývojáři používají Cursor, Claude Code, Codex CLI, OpenClaw, Gemini CLI, K
- **Průvodce zaváděním** 4krokové nastavení pro začínající uživatele
- **Jeden koncový bod, všechny modely** jednou nakonfigurujte `http://localhost:20128/v1` a získejte přístup k více než 44 poskytovatelům
<details>
<summary><b>🔑 8. „Správa OAuth tokenů od více poskytovatelů je peklo“</b></summary>
</details>
@@ -288,9 +263,6 @@ Claude Code, Codex, Gemini CLI, Copilot všechny používají OAuth 2.0 s to
- **OAuth Behind Nginx** — Používá `window.location.origin` pro kompatibilitu s reverzní proxy
- **Průvodce vzdáleným OAuth** Podrobný návod k přihlašovacím údajům Google Cloud na VPS/Dockeru
<details>
<summary><b>📊 9. „Nevím, kolik utrácím ani kde“</b></summary>
</details>
@@ -305,9 +277,6 @@ Vývojáři používají více placených poskytovatelů, ale nemají jednotný
- **Statistiky použití pro každý klíč API** — Počet požadavků a časové razítko posledního použití pro každý klíč
- **Analytický panel** Statistické karty, graf využití modelu, tabulka poskytovatelů s mírou úspěšnosti a latencí
<details>
<summary><b>🐛 10. „Nedokážu diagnostikovat chyby a problémy ve volání umělé inteligence.“</b></summary>
</details>
@@ -324,9 +293,6 @@ Když volání selže, vývojář neví, zda se jednalo o limit rychlosti, vypr
- **Souborové protokolování s rotací** Konzolový interceptor zachycuje vše do protokolu JSON s rotací na základě velikosti
- **Zpráva o systémových informacích** — příkaz `npm run system-info` vygeneruje `system-info.txt` s kompletním popisem vašeho prostředí (verze uzlu, verze OmniRoute, operační systém, nástroje CLI, stav Dockeru/PM2). Přiložte jej při hlášení problémů pro okamžité třídění.
<details>
<summary><b>🏗️ 11. „Nasazení a údržba brány je složitá“</b></summary>
</details>
@@ -343,9 +309,6 @@ Instalace, konfigurace a údržba AI proxy v různých prostředích (lokální,
- **Cloud Sync** Konfigurace synchronizace mezi zařízeními pomocí Cloudflare Workers
- **Zálohy databází** — Automatické zálohování, obnovení, export a import všech nastavení
<details>
<summary><b>🌍 12. „Rozhraní je pouze v angličtině a můj tým nemluví anglicky“</b></summary>
</details>
@@ -359,9 +322,6 @@ Týmy v neanglicky mluvících zemích, zejména v Latinské Americe, Asii a Evr
- **Vícejazyčné soubory README** — 30 kompletních překladů dokumentace
- **Výběr jazyka** — Ikona glóbu v záhlaví pro přepínání v reálném čase
<details>
<summary><b>🔄 13. „Potřebuji víc než jen chat potřebuji vložené soubory, obrázky, zvuk.“</b></summary>
</details>
@@ -380,9 +340,6 @@ Umělá inteligence není jen dokončování chatu. Vývojáři potřebují gene
- **Změna pořadí**`/v1/rerank` — Změna pořadí relevance dokumentu
- **Responses API** — Plná podpora `/v1/responses` pro Codex
<details>
<summary><b>🧪 14. „Nemám způsob, jak testovat a porovnávat kvalitu napříč modely.“</b></summary>
</details>
@@ -397,9 +354,6 @@ Vývojáři chtějí vědět, který model je pro jejich případ použití nejl
- **Tester chatu** — Kompletní okružní cesta s vizuálním vykreslováním odpovědí
- **Živý monitor** — Stream všech požadavků procházejících proxy serverem v reálném čase
<details>
<summary><b>📈 15. „Potřebuji škálovat bez ztráty výkonu“</b></summary>
</details>
@@ -415,9 +369,6 @@ S rostoucím objemem požadavků generují stejné otázky bez ukládání do me
- **Mezipaměť pro ověření klíčů API** — třívrstvá mezipaměť pro výkon produkčního prostředí
- **Dashboard s telemetrií** latence p50/p95/p99, statistiky mezipaměti, dostupnost
<details>
<summary><b>🤖 16. „Chci mít chování modelů globálně pod kontrolou“</b></summary>
</details>
@@ -434,9 +385,6 @@ Vývojáři, kteří chtějí všechny odpovědi v určitém jazyce, se specific
- **Přepínání poskytovatele** Povolení/zakázání všech připojení pro poskytovatele jedním kliknutím
- **Blokovaní poskytovatelé** Vyloučení konkrétních poskytovatelů ze seznamu `/v1/models`
<details>
<summary><b>🧰 17. „Potřebuji nástroje MCP jako prvotřídní produktové funkce.“</b></summary>
</details>
@@ -449,9 +397,6 @@ Mnoho bran umělé inteligence odhaluje MCP pouze jako skrytý implementační d
- Vyhrazená stránka pro správu MCP s procesy, nástroji, rozsahy a auditem
- Vestavěný rychlý start pro `omniroute --mcp` a onboarding klienta
<details>
<summary><b>🧠 18. „Potřebuji orchestraci A2A se synchronizací a cestami úloh streamu.“</b></summary>
</details>
@@ -464,9 +409,6 @@ Pracovní postupy agentů vyžadují jak přímé odpovědi, tak dlouhodobé str
- Streamování SSE s šířením stavu terminálu
- Rozhraní API životního cyklu úloh pro `tasks/get` a `tasks/cancel`
<details>
<summary><b>🛰️ 19. „Potřebuji skutečný stav procesu MCP, ne odhadovaný stav.“</b></summary>
</details>
@@ -479,9 +421,6 @@ Provozní týmy potřebují vědět, zda je MCP skutečně aktivní, nejen zda j
- API stavu MCP kombinující prezenční signál a nedávnou aktivitu
- Karty stavu uživatelského rozhraní pro zobrazení aktuálnosti procesů/provozuschopnosti/prezenčního signálu
<details>
<summary><b>📋 20. „Potřebuji auditovatelné provedení nástroje MCP“</b></summary>
</details>
@@ -494,9 +433,6 @@ Když nástroje mění konfiguraci nebo spouštějí operační akce, týmy pot
- Filtruje podle nástroje, úspěchu/neúspěchu, klíče API a stránkování
- Tabulka auditu dashboardu + koncové body statistik pro automatizaci
<details>
<summary><b>🔐 21. „Potřebuji omezená oprávnění MCP pro každou integraci.“</b></summary>
</details>
@@ -509,9 +445,6 @@ Různí klienti by měli mít přístup ke kategoriím nástrojů s nejnižším
- Vynucení rozsahu a viditelnost v uživatelském rozhraní správy MCP
- Bezpečná výchozí poloha pro provozní nástroje
<details>
<summary><b>⚙️ 22. „Potřebuji provozní kontroly bez nutnosti přesouvání“</b></summary>
</details>
@@ -524,9 +457,6 @@ Týmy potřebují rychlé změny v běhovém prostředí během incidentů nebo
- Používejte profily odolnosti z předdefinovaných balíčků zásad
- Resetujte stav jističe ze stejného ovládacího panelu
<details>
<summary><b>🔄 23. „Potřebuji živý přehled o životním cyklu úkolů A2A a jejich zrušení.“</b></summary>
</details>
@@ -539,9 +469,6 @@ Bez přehledu o životním cyklu je obtížné třídit incidenty úkolů.
- Podrobný přehled metadat úloh, událostí a artefaktů
- Koncový bod zrušení úlohy a akce uživatelského rozhraní s potvrzením
<details>
<summary><b>🌊 24. „Potřebuji metriky aktivního streamu pro A2A zátěž“</b></summary>
</details>
@@ -554,9 +481,6 @@ Streamovací pracovní postupy vyžadují provozní přehled o souběžnosti a
- Časové razítko posledního úkolu a počty pro jednotlivé stavy
- Karty A2A dashboardu pro monitorování provozu v reálném čase
<details>
<summary><b>🪪 25. „Potřebuji standardní vyhledávání agentů pro klienty“</b></summary>
</details>
@@ -569,9 +493,6 @@ Externí klienti a orchestratoři potřebují pro onboarding strojově čitelná
- Schopnosti a dovednosti zobrazené v uživatelském rozhraní pro správu
- API pro stav A2A zahrnuje metadata pro zjišťování pro automatizaci
<details>
<summary><b>🧭 26. „Potřebuji v uživatelském rozhraní produktu zjistitelnost protokolu.“</b></summary>
</details>
@@ -584,9 +505,6 @@ Pokud uživatelé nemohou objevit protokolové povrchy, kvalita přijetí a podp
- Přepínání stavu inline služby (Online/Offline) pro MCP a A2A
- Odkazy z přehledu na vyhrazené karty pro správu
<details>
<summary><b>🧪 27. „Potřebuji komplexní ověření protokolu se skutečnými klienty.“</b></summary>
</details>
@@ -599,9 +517,6 @@ Simulované testy nestačí k ověření kompatibility protokolu před vydáním
- Klientské testy A2A pro toky zjišťování, odesílání, streamování, načítání a zrušení
- Křížová kontrola tvrzení oproti API pro audit MCP a úkoly A2A
<details>
<summary><b>📡 28. „Potřebuji jednotnou pozorovatelnost napříč všemi rozhraními“</b></summary>
</details>
@@ -614,9 +529,6 @@ Rozdělení pozorovatelnosti podle protokolu vytváří slepá místa a delší
- Stav + audit + telemetrie požadavků napříč vrstvami OpenAI, MCP a A2A
- Provozní API pro stav a automatizaci
<details>
<summary><b>💼 29. „Potřebuji jeden runtime pro proxy + nástroje + orchestraci agentů“</b></summary>
</details>
@@ -629,9 +541,6 @@ Spouštění mnoha samostatných služeb zvyšuje provozní náklady a počet po
- Sdílené ověřování, odolnost, úložiště dat a pozorovatelnost
- Konzistentní model politik napříč všemi interakčními plochami
<details>
<summary><b>🚀 30. „Potřebuji agentské pracovní postupy bez slepení kódu.“</b></summary>
</details>
@@ -644,9 +553,6 @@ Týmy ztrácejí rychlost při spojování více ad-hoc služeb a skriptů.
- Vestavěná uživatelská rozhraní pro správu protokolů a cesty pro ověřování kouře
- Základy připravené pro produkční prostředí (zabezpečení, protokolování, odolnost, zálohování)
### Příklady herních plánů (integrované případy užití)
**Příručka A: Maximalizace placeného předplatného + levné zálohování**
@@ -701,13 +607,13 @@ Outcome: deep fallback depth for deadline-critical workloads
> Nastavte si kódování s umělou inteligencí během několika minut za **0 $/měsíc** . Propojte tyto bezplatné účty a využijte vestavěnou kombinaci **Free Stack** .
Krok | Akce | Poskytovatelé odemčeni
--- | --- | ---
1 | Připojení **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 **neomezeně**
2 | Připojení k **iFlow** (Google OAuth) | kimi-k2-myšlení, qwen3-coder-plus, deepseek-r1... — **neomezeně**
3 | Připojení **Qwen** (kód zařízení) | qwen3-coder-plus, qwen3-coder-flash... — **neomezeně**
4 | Připojení **rozhraní příkazového řádku Gemini** (Google OAuth) | gemini-3-flash, gemini-2.5-pro — **180 000 GBP/měsíc zdarma**
5 | `/dashboard/combos` → Šablona **Free Stack (0 $)** | Automatické zařazení všech bezplatných poskytovatelů do routingu
| Krok | Akce | Poskytovatelé odemčeni |
| ---- | -------------------------------------------------------------- | ----------------------------------------------------------------- |
| 1 | Připojení **Kiro** (AWS Builder ID OAuth) | Claude Sonnet 4.5, Haiku 4.5 **neomezeně** |
| 2 | Připojení k **iFlow** (Google OAuth) | kimi-k2-myšlení, qwen3-coder-plus, deepseek-r1... — **neomezeně** |
| 3 | Připojení **Qwen** (kód zařízení) | qwen3-coder-plus, qwen3-coder-flash... — **neomezeně** |
| 4 | Připojení **rozhraní příkazového řádku Gemini** (Google OAuth) | gemini-3-flash, gemini-2.5-pro — **180 000 GBP/měsíc zdarma** |
| 5 | `/dashboard/combos` → Šablona **Free Stack (0 $)** | Automatické zařazení všech bezplatných poskytovatelů do routingu |
**V libovolném IDE/CLI naveďte:** `http://localhost:20128/v1` · Klíč API: `any-string` · Hotovo.
@@ -732,13 +638,13 @@ omniroute
Dashboard se otevírá na `http://localhost:20128` a základní URL API je `http://localhost:20128/v1` .
Příkaz | Popis
--- | ---
`omniroute` | Spuštění serveru ( `PORT=20128` , API a dashboard na stejném portu)
`omniroute --port 3000` | Nastavte kanonický/API port na 3000
`omniroute --mcp` | Spuštění MCP serveru (transport stdio)
`omniroute --no-open` | Neotevírat prohlížeč automaticky
`omniroute --help` | Zobrazit nápovědu
| Příkaz | Popis |
| ----------------------- | ------------------------------------------------------------------- |
| `omniroute` | Spuštění serveru ( `PORT=20128` , API a dashboard na stejném portu) |
| `omniroute --port 3000` | Nastavte kanonický/API port na 3000 |
| `omniroute --mcp` | Spuštění MCP serveru (transport stdio) |
| `omniroute --no-open` | Neotevírat prohlížeč automaticky |
| `omniroute --help` | Zobrazit nápovědu |
Volitelný režim s rozděleným portem:
@@ -847,10 +753,10 @@ docker compose --profile base up -d
docker compose --profile cli up -d
```
Obraz | Štítek | Velikost | Popis
--- | --- | --- | ---
`diegosouzapw/omniroute` | `latest` | ~250 MB | Nejnovější stabilní verze
`diegosouzapw/omniroute` | `1.0.3` | ~250 MB | Aktuální verze
| Obraz | Štítek | Velikost | Popis |
| ------------------------ | -------- | -------- | ------------------------- |
| `diegosouzapw/omniroute` | `latest` | ~250 MB | Nejnovější stabilní verze |
| `diegosouzapw/omniroute` | `1.0.3` | ~250 MB | Aktuální verze |
---
@@ -893,41 +799,47 @@ Po minimalizaci se OmniRoute nachází v systémové liště a nabízí rychlé
## 💰 Přehled cen
Úroveň | Poskytovatel | Náklady | Obnovení kvóty | Nejlepší pro
--- | --- | --- | --- | ---
**💳 PŘEDPLATNÉ** | Claude Code (profesionál) | 20 dolarů měsíčně | 5 hodin + týdně | Již přihlášen/a k odběru
| Kodex (Plus/Pro) | 20200 USD/měsíc | 5 hodin + týdně | Uživatelé OpenAI
| Rozhraní příkazového řádku Gemini | **UVOLNIT** | 180 tisíc měsíčně + 1 tisíc denně | Každý!
| GitHub Copilot | 1019 USD/měsíc | Měsíční | Uživatelé GitHubu
**🔑 KLÍČ API** | NVIDIA NIM | **ZDARMA** (vývoj navždy) | ~40 ot./min | 70+ otevřených modelů
| Mozky | **ZDARMA** (1 milion tok/den) | 60 000 otáček za minutu / 30 ot./min | Nejrychlejší na světě
| Groq | **ZDARMA** (30 ot./min.) | 14,4 tisíc otáček za minutu | Ultrarychlá lama/gema
| DeepSeek V3.2 | 0,27/1,10 USD za 1 milion | Žádný | Nejlepší zdůvodnění ceny a kvality
| xAI Grok-4 Rychlý | **0,20/0,50 USD za 1 milion** 🆕 | Žádný | Nejrychlejší + volání nástroje, ultranízké
| xAI Grok-4 (standardní) | 0,20/1,50 USD za 1 milion 🆕 | Žádný | Vlajková loď Reasoning od xAI
| Mistral | Zkušební verze zdarma + placené | Omezená sazba | Evropská umělá inteligence
| OpenRouter | Platba za použití | Žádný | Více než 100 modelů agregováno.
**💰 LEVNÉ** | GLM-5 (přes Z.AI) 🆕 | 0,5 USD/1 milion | Denně v 10:00 | Výstup 128 tisíc obrazových bodů, nejnovější vlajková loď
| GLM-4.7 | 0,6 USD/1 milion | Denně v 10:00 | Záloha rozpočtu
| MiniMax M2.5 🆕 | Vstup 0,3 USD/1 milion | 5hodinové válcování | Úvaha + agentní úkoly
| MiniMax M2.1 | 0,2 USD/1 milion | 5hodinové válcování | Nejlevnější varianta
| Kimi K2.5 (Moonshot API) 🆕 | Platba za použití | Žádný | Přímý přístup k Moonshot API
| Kimi K2 | 9 dolarů měsíčně bez závazků | 10 milionů tokenů/měsíc | Předvídatelné náklady
**🆓 ZDARMA** | iFlow | **0 dolarů** | Neomezený | 5 modelů neomezeně
| Qwen | **0 dolarů** | Neomezený | 4 modely neomezeně
| Kiro | **0 dolarů** | Neomezený | Claude Sonnet/Haiku (tvorce AWS)
| Úroveň | Poskytovatel | Náklady | Obnovení kvóty | Nejlepší pro |
| --------------------------------- | -------------------------------- | ------------------------------------ | ------------------------------------------ | --------------------------------------------------------- |
| **💳 PŘEDPLATNÉ** | Claude Code (profesionál) | 20 dolarů měsíčně | 5 hodin + týdně | Již přihlášen/a k odběru |
| Kodex (Plus/Pro) | 20200 USD/měsíc | 5 hodin + týdně | Uživatelé OpenAI |
| Rozhraní příkazového řádku Gemini | **UVOLNIT** | 180 tisíc měsíčně + 1 tisíc denně | Každý! |
| GitHub Copilot | 1019 USD/měsíc | Měsíční | Uživatelé GitHubu |
| **🔑 KLÍČ API** | NVIDIA NIM | **ZDARMA** (vývoj navždy) | ~40 ot./min | 70+ otevřených modelů |
| Mozky | **ZDARMA** (1 milion tok/den) | 60 000 otáček za minutu / 30 ot./min | Nejrychlejší na světě |
| Groq | **ZDARMA** (30 ot./min.) | 14,4 tisíc otáček za minutu | Ultrarychlá lama/gema |
| DeepSeek V3.2 | 0,27/1,10 USD za 1 milion | Žádný | Nejlepší zdůvodnění ceny a kvality |
| xAI Grok-4 Rychlý | **0,20/0,50 USD za 1 milion** 🆕 | Žádný | Nejrychlejší + volání nástroje, ultranízké |
| xAI Grok-4 (standardní) | 0,20/1,50 USD za 1 milion 🆕 | Žádný | Vlajková loď Reasoning od xAI |
| Mistral | Zkušební verze zdarma + placené | Omezená sazba | Evropská umělá inteligence |
| OpenRouter | Platba za použití | Žádný | Více než 100 modelů agregováno. |
| **💰 LEVNÉ** | GLM-5 (přes Z.AI) 🆕 | 0,5 USD/1 milion | Denně v 10:00 | Výstup 128 tisíc obrazových bodů, nejnovější vlajková loď |
| GLM-4.7 | 0,6 USD/1 milion | Denně v 10:00 | Záloha rozpočtu |
| MiniMax M2.5 🆕 | Vstup 0,3 USD/1 milion | 5hodinové válcování | Úvaha + agentní úkoly |
| MiniMax M2.1 | 0,2 USD/1 milion | 5hodinové válcování | Nejlevnější varianta |
| Kimi K2.5 (Moonshot API) 🆕 | Platba za použití | Žádný | Přímý přístup k Moonshot API |
| Kimi K2 | 9 dolarů měsíčně bez závazků | 10 milionů tokenů/měsíc | Předvídatelné náklady |
| **🆓 ZDARMA** | iFlow | **0 dolarů** | Neomezený | 5 modelů neomezeně |
| Qwen | **0 dolarů** | Neomezený | 4 modely neomezeně |
| Kiro | **0 dolarů** | Neomezený | Claude Sonnet/Haiku (tvorce AWS) |
> 🆕 **Přidány nové modely (březen 2026):** řada Grok-4 Fast za 0,20 USD/0,50 USD/M (benchmarkováno na 1143 ms o 30 % rychlejší než Gemini 2.5 Flash), GLM-5 přes Z.AI s výstupem 128K, uvažování MiniMax M2.5, aktualizované ceny DeepSeek V3.2, Kimi K2.5 přes Moonshot Direct API.
**💡 Kombinovaný balík za 0 $ — Kompletní bezplatná instalace:**
```
Gemini CLI (180K/mo free)
→ iFlow (unlimited: kimi-k2-thinking, qwen3-coder-plus, deepseek-r1)
→ Kiro (Claude Sonnet 4.5 + Haiku — unlimited, via AWS Builder ID)
→ Qwen (4 models — unlimited)
→ Groq (14.4K req/day — ultra-fast)
→ NVIDIA NIM (70+ models — 40 RPM forever)
# 🆓 Ultimate Free Stack 2026 — 11 Providers, $0 Forever
Kiro (kr/) → Claude Sonnet/Haiku UNLIMITED
iFlow (if/) → kimi-k2-thinking, qwen3-coder-plus, deepseek-r1 UNLIMITED
LongCat Lite (lc/) → LongCat-Flash-Lite — 50M tokens/day 🔥
Pollinations (pol/) → GPT-5, Claude, DeepSeek, Llama 4 — no key needed
Qwen (qw/) → qwen3-coder-plus, qwen3-coder-flash, qwen3-coder-next UNLIMITED
Gemini (gemini/) → Gemini 2.5 Flash — 1,500 req/day free API key
Cloudflare AI (cf/) → Llama 70B, Gemma 3, Mistral — 10K Neurons/day
Scaleway (scw/) → Qwen3 235B, Llama 70B — 1M free tokens (EU)
Groq (groq/) → Llama/Gemma ultra-fast — 14.4K req/day
NVIDIA NIM (nvidia/) → 70+ open models — 40 RPM forever
Cerebras (cerebras/) → Llama/Qwen world-fastest — 1M tok/day
```
**Nulové náklady. Nikdy nepřestávejte s kódováním.** Nakonfigurujte si to jako jednu kombinaci OmniRoute a všechny záložní režimy se provede automaticky žádné ruční přepínání.
@@ -942,59 +854,59 @@ Gemini CLI (180K/mo free)
### 🔵 CLAUDE MODELS (přes Kiro — AWS Builder ID)
Model | Předpona | Omezit | Limit rychlosti
--- | --- | --- | ---
`claude-sonnet-4.5` | `kr/` | **Neomezený** | Žádný hlášený denní limit
`claude-haiku-4.5` | `kr/` | **Neomezený** | Žádný hlášený denní limit
`claude-opus-4.6` | `kr/` | **Neomezený** | Nejnovější opus od Kira
| Model | Předpona | Omezit | Limit rychlosti |
| ------------------- | -------- | ------------- | ------------------------- |
| `claude-sonnet-4.5` | `kr/` | **Neomezený** | Žádný hlášený denní limit |
| `claude-haiku-4.5` | `kr/` | **Neomezený** | Žádný hlášený denní limit |
| `claude-opus-4.6` | `kr/` | **Neomezený** | Nejnovější opus od Kira |
### 🟢 MODELY IFLOW (Bezplatné OAuth — bez nutnosti platit kreditní kartou)
Model | Předpona | Omezit | Limit rychlosti
--- | --- | --- | ---
`kimi-k2-thinking` | `if/` | **Neomezený** | Žádný hlášený strop
`qwen3-coder-plus` | `if/` | **Neomezený** | Žádný hlášený strop
`deepseek-r1` | `if/` | **Neomezený** | Žádný hlášený strop
`minimax-m2.1` | `if/` | **Neomezený** | Žádný hlášený strop
`kimi-k2` | `if/` | **Neomezený** | Žádný hlášený strop
| Model | Předpona | Omezit | Limit rychlosti |
| ------------------ | -------- | ------------- | ------------------- |
| `kimi-k2-thinking` | `if/` | **Neomezený** | Žádný hlášený strop |
| `qwen3-coder-plus` | `if/` | **Neomezený** | Žádný hlášený strop |
| `deepseek-r1` | `if/` | **Neomezený** | Žádný hlášený strop |
| `minimax-m2.1` | `if/` | **Neomezený** | Žádný hlášený strop |
| `kimi-k2` | `if/` | **Neomezený** | Žádný hlášený strop |
### 🟡 MODELY QWEN (Ověření kódu zařízení)
Model | Předpona | Omezit | Limit rychlosti
--- | --- | --- | ---
`qwen3-coder-plus` | `qw/` | **Neomezený** | Žádný hlášený strop
`qwen3-coder-flash` | `qw/` | **Neomezený** | Žádný hlášený strop
`qwen3-coder-next` | `qw/` | **Neomezený** | Žádný hlášený strop
`vision-model` | `qw/` | **Neomezený** | Multimodální (obrázky)
| Model | Předpona | Omezit | Limit rychlosti |
| ------------------- | -------- | ------------- | ---------------------- |
| `qwen3-coder-plus` | `qw/` | **Neomezený** | Žádný hlášený strop |
| `qwen3-coder-flash` | `qw/` | **Neomezený** | Žádný hlášený strop |
| `qwen3-coder-next` | `qw/` | **Neomezený** | Žádný hlášený strop |
| `vision-model` | `qw/` | **Neomezený** | Multimodální (obrázky) |
### 🟣 Rozhraní GEMINI CLI (Google OAuth)
Model | Předpona | Omezit | Limit rychlosti
--- | --- | --- | ---
`gemini-3-flash-preview` | `gc/` | **180 tisíc tok/měsíc** + 1 tisíc/den | Měsíční reset
`gemini-2.5-pro` | `gc/` | 180 tisíc měsíčně (sdílený bazén) | Vysoká kvalita
| Model | Předpona | Omezit | Limit rychlosti |
| ------------------------ | -------- | ------------------------------------- | --------------- |
| `gemini-3-flash-preview` | `gc/` | **180 tisíc tok/měsíc** + 1 tisíc/den | Měsíční reset |
| `gemini-2.5-pro` | `gc/` | 180 tisíc měsíčně (sdílený bazén) | Vysoká kvalita |
### ⚫ NVIDIA NIM (Bezplatný klíč API — build.nvidia.com)
Úroveň | Denní limit | Limit rychlosti | Poznámky
--- | --- | --- | ---
Zdarma (vývojář) | Žádný limit tokenů | **~40 ot./min** | Více než 70 modelů; přechod na čisté limity sazeb v polovině roku 2025
| Úroveň | Denní limit | Limit rychlosti | Poznámky |
| ---------------- | ------------------ | --------------- | ---------------------------------------------------------------------- |
| Zdarma (vývojář) | Žádný limit tokenů | **~40 ot./min** | Více než 70 modelů; přechod na čisté limity sazeb v polovině roku 2025 |
Oblíbené bezplatné modely: `moonshotai/kimi-k2.5` (Kimi K2.5), `z-ai/glm4.7` (GLM 4.7), `deepseek-ai/deepseek-v3.2` (DeepSeek V3.2), `nvidia/llama-3.3-70b-instruct` , `deepseek/deepseek-r1`
### ⚪ CEREBRAS (Bezplatný klíč API — inference.cerebras.ai)
Úroveň | Denní limit | Limit rychlosti | Poznámky
--- | --- | --- | ---
Uvolnit | **1 milion tokenů/den** | 60 000 otáček za minutu / 30 ot./min | Nejrychlejší inference LLM na světě; denně se resetuje
| Úroveň | Denní limit | Limit rychlosti | Poznámky |
| ------- | ----------------------- | ------------------------------------ | ------------------------------------------------------ |
| Uvolnit | **1 milion tokenů/den** | 60 000 otáček za minutu / 30 ot./min | Nejrychlejší inference LLM na světě; denně se resetuje |
Dostupné zdarma: `llama-3.3-70b` , `llama-3.1-8b` , `deepseek-r1-distill-llama-70b`
### 🔴 GROQ (Bezplatný API klíč — console.groq.com)
Úroveň | Denní limit | Limit rychlosti | Poznámky
--- | --- | --- | ---
Uvolnit | **14,4 tisíc otáček za minutu** | 30 ot./min na model | Žádná kreditní karta; limit 429, neúčtováno
| Úroveň | Denní limit | Limit rychlosti | Poznámky |
| ------- | ------------------------------- | ------------------- | ------------------------------------------- |
| Uvolnit | **14,4 tisíc otáček za minutu** | 30 ot./min na model | Žádná kreditní karta; limit 429, neúčtováno |
K dispozici zdarma: `llama-3.3-70b-versatile` , `gemma2-9b-it` , `mixtral-8x7b` , `whisper-large-v3`
@@ -1016,11 +928,11 @@ K dispozici zdarma: `llama-3.3-70b-versatile` , `gemma2-9b-it` , `mixtral-8x7b`
> Přepisujte libovolné audio/video za **0 $** Deepgram leady za 200 $ zdarma, AssemblyAI za 50 $ jako záložní nástroj, Groq Whisper jako neomezená nouzová záloha.
Poskytovatel | Bezplatné kredity | Nejlepší model | Limit rychlosti
--- | --- | --- | ---
🟢 **Deepgram** | **200 dolarů zdarma** (registrace) | `nova-3` — nejvyšší přesnost, více než 30 jazyků | Žádný limit RPM pro kredity zdarma
🔵 **AssemblyAI** | **50 dolarů zdarma** (registrace) | `universal-3-pro` — kapitoly, sentiment, osobní údaje | Žádný limit RPM pro kredity zdarma
🔴 **Groq** | **Navždy zdarma** | `whisper-large-v3` — OpenAI Šepot | 30 ot./min (omezená rychlost)
| Poskytovatel | Bezplatné kredity | Nejlepší model | Limit rychlosti |
| ----------------- | ---------------------------------- | ----------------------------------------------------- | ---------------------------------- |
| 🟢 **Deepgram** | **200 dolarů zdarma** (registrace) | `nova-3` — nejvyšší přesnost, více než 30 jazyků | Žádný limit RPM pro kredity zdarma |
| 🔵 **AssemblyAI** | **50 dolarů zdarma** (registrace) | `universal-3-pro` — kapitoly, sentiment, osobní údaje | Žádný limit RPM pro kredity zdarma |
| 🔴 **Groq** | **Navždy zdarma** | `whisper-large-v3` — OpenAI Šepot | 30 ot./min (omezená rychlost) |
**Navrhovaná kombinace v `/dashboard/combos` :**
@@ -1041,118 +953,118 @@ OmniRoute v2.0 je navržen jako operační platforma, nikoli pouze jako proxy pr
### 🆕 Nové — Vylepšení inspirovaná ClawRouterem (březen 2026)
Funkce | Co to dělá
--- | ---
**Grok-4 Rychlá rodina** | Modely xAI za 0,20 USD/0,50 USD/M v benchmarku 1143 ms (o 30 % rychlejší než Gemini 2.5 Flash)
🧠 **GLM-5 přes Z.AI** | 128 tisíc výstupních dat, 0,5 USD/1 milion USD nejnovější vlajková loď rodiny GLM
🔮 **MiniMax M2.5** | Úvaha + agentní úkoly za 0,30 USD/1 milion významný upgrade oproti M2.1
🎯 **Příznak volání nástroje pro každý model** | `toolCalling: true/false` v registru — AutoCombo přeskakuje modely, které nepodporují nástroje.
🌍 **Detekce vícejazyčného záměru** | Klíčová slova PT/ZH/ES/AR v bodování AutoCombo lepší výběr modelu pro neanglický obsah
📊 **Záložní metody řízené benchmarkem** | Skutečná latence p95 z živých požadavků poskytuje kombinované skóre AutoCombo se učí ze skutečných dat
🔁 **Požádat o deduplikaci** | Okno pro deduplikaci na základě hashování obsahu bezpečné pro více agentů, zabraňuje duplicitním platbám
🔌 **Strategie pro zásuvné routery** | Rozšiřitelné rozhraní `RouterStrategy` přidejte si vlastní logiku směrování jako pluginy
| Funkce | Co to dělá |
| ---------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| **Grok-4 Rychlá rodina** | Modely xAI za 0,20 USD/0,50 USD/M v benchmarku 1143 ms (o 30 % rychlejší než Gemini 2.5 Flash) |
| 🧠 **GLM-5 přes Z.AI** | 128 tisíc výstupních dat, 0,5 USD/1 milion USD nejnovější vlajková loď rodiny GLM |
| 🔮 **MiniMax M2.5** | Úvaha + agentní úkoly za 0,30 USD/1 milion významný upgrade oproti M2.1 |
| 🎯 **Příznak volání nástroje pro každý model** | `toolCalling: true/false` v registru — AutoCombo přeskakuje modely, které nepodporují nástroje. |
| 🌍 **Detekce vícejazyčného záměru** | Klíčová slova PT/ZH/ES/AR v bodování AutoCombo lepší výběr modelu pro neanglický obsah |
| 📊 **Záložní metody řízené benchmarkem** | Skutečná latence p95 z živých požadavků poskytuje kombinované skóre AutoCombo se učí ze skutečných dat |
| 🔁 **Požádat o deduplikaci** | Okno pro deduplikaci na základě hashování obsahu bezpečné pro více agentů, zabraňuje duplicitním platbám |
| 🔌 **Strategie pro zásuvné routery** | Rozšiřitelné rozhraní `RouterStrategy` přidejte si vlastní logiku směrování jako pluginy |
### 🚀 Předchozí verze v2.0.9+ — Hřiště, otisky prstů v CLI a ACP
Funkce | Co to dělá
--- | ---
🎮 **Modelové hřiště** | Stránka řídicího panelu pro přímé testování libovolného modelu selektory poskytovatele/modelu/koncového bodu, editor Monaco, streamování, přerušení, načasování
🔏 **Porovnávání otisků prstů v CLI** | Řazení hlaviček/těl serveru podle poskytovatele tak, aby odpovídalo nativním podpisům CLI přepínání pro jednotlivé poskytovatele v Nastavení &gt; Zabezpečení. **Vaše IP adresa proxy serveru je zachována.**
🤝 **Podpora ACP (Agent Client Protocol)** | Vyhledávání agentů CLI (Codex, Claude, Goose, Gemini CLI, OpenClaw + 9 dalších), generátor procesů, koncový bod `/api/acp/agents`
🤖 **Řídicí panel agentů ACP** | Ladění Stránka Agenti — mřížka 14 agentů se stavem instalace, verzí a formulářem pro vlastní agenta pro libovolný nástroj CLI. Uživatelé **OpenCode** získají tlačítko „Stáhnout opencode.json“, které automaticky vygeneruje konfiguraci připravenou k použití se všemi dostupnými modely.
🔧 **Směrování `apiFormat` pro vlastní model** | Vlastní modely s `apiFormat: "responses"` nyní správně směrují do překladače Responses API.
🏢 **Izolace pracovního prostoru Codexu** | Více pracovních prostorů Codexu na jeden e-mail OAuth správně odděluje připojení podle ID pracovního prostoru
🔄 **Automatická aktualizace elektronů** | Desktopová aplikace kontroluje aktualizace + automaticky se instaluje po restartu
| Funkce | Co to dělá |
| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🎮 **Modelové hřiště** | Stránka řídicího panelu pro přímé testování libovolného modelu selektory poskytovatele/modelu/koncového bodu, editor Monaco, streamování, přerušení, načasování |
| 🔏 **Porovnávání otisků prstů v CLI** | Řazení hlaviček/těl serveru podle poskytovatele tak, aby odpovídalo nativním podpisům CLI přepínání pro jednotlivé poskytovatele v Nastavení &gt; Zabezpečení. **Vaše IP adresa proxy serveru je zachována.** |
| 🤝 **Podpora ACP (Agent Client Protocol)** | Vyhledávání agentů CLI (Codex, Claude, Goose, Gemini CLI, OpenClaw + 9 dalších), generátor procesů, koncový bod `/api/acp/agents` |
| 🤖 **Řídicí panel agentů ACP** | Ladění Stránka Agenti — mřížka 14 agentů se stavem instalace, verzí a formulářem pro vlastní agenta pro libovolný nástroj CLI. Uživatelé **OpenCode** získají tlačítko „Stáhnout opencode.json“, které automaticky vygeneruje konfiguraci připravenou k použití se všemi dostupnými modely. |
| 🔧 **Směrování `apiFormat` pro vlastní model** | Vlastní modely s `apiFormat: "responses"` nyní správně směrují do překladače Responses API. |
| 🏢 **Izolace pracovního prostoru Codexu** | Více pracovních prostorů Codexu na jeden e-mail OAuth správně odděluje připojení podle ID pracovního prostoru |
| 🔄 **Automatická aktualizace elektronů** | Desktopová aplikace kontroluje aktualizace + automaticky se instaluje po restartu |
### 🤖 Operace s agenty a protokoly (v2.0)
Funkce | Co to dělá
--- | ---
🔧 **MCP Server (16 nástrojů)** | Nástroje IDE/agent prostřednictvím 3 transportů: stdio, SSE ( `/api/mcp/sse` ), Streamovatelný HTTP ( `/api/mcp/stream` )
🤝 **A2A server (JSON-RPC + SSE)** | Spouštění úloh mezi agenty se synchronizací a streamováním
🧭 **Konsolidovaná stránka koncových bodů** | Stránka pro správu s kartami Endpoint Proxy, MCP, A2A a API Endpoints
🎚️ **Přepínače pro povolení/zakázání služby** | Přepínače ZAP/VYP pro MCP a A2A s trvalým nastavením (výchozí: VYP)
🛰️ **Srdeční tep za běhu MCP** | Skutečný stav procesu (pid, doba provozuschopnosti, stáří heartbeatu, transport, režim rozsahu)
📋 **Auditní záznam MCP** | Filtrovatelné protokoly auditu s hodnocením úspěchu/neúspěchu a klíčovým přiřazením
🔐 **Vynucování rozsahu MCP** | 9 podrobných oprávnění pro řízený přístup k nástrojům
📡 **Správa životního cyklu úkolů A2A** | Seznam/filtrování úloh, kontrola událostí/artefaktů, zrušení spuštěných úloh
📋 **Objevení karty agenta** | `/.well-known/agent.json` pro automatické vyhledávání klientů
🧪 **Testovací postroj Protocol E2E** | Skutečné MCP SDK + toky klientů A2A v `test:protocols:e2e`
⚙️ **Provozní kontroly** | Kombinace přepínačů, použití profilů odolnosti, resetování jističů z jednoho ovládacího panelu
| Funkce | Co to dělá |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- |
| 🔧 **MCP Server (16 nástrojů)** | Nástroje IDE/agent prostřednictvím 3 transportů: stdio, SSE ( `/api/mcp/sse` ), Streamovatelný HTTP ( `/api/mcp/stream` ) |
| 🤝 **A2A server (JSON-RPC + SSE)** | Spouštění úloh mezi agenty se synchronizací a streamováním |
| 🧭 **Konsolidovaná stránka koncových bodů** | Stránka pro správu s kartami Endpoint Proxy, MCP, A2A a API Endpoints |
| 🎚️ **Přepínače pro povolení/zakázání služby** | Přepínače ZAP/VYP pro MCP a A2A s trvalým nastavením (výchozí: VYP) |
| 🛰️ **Srdeční tep za běhu MCP** | Skutečný stav procesu (pid, doba provozuschopnosti, stáří heartbeatu, transport, režim rozsahu) |
| 📋 **Auditní záznam MCP** | Filtrovatelné protokoly auditu s hodnocením úspěchu/neúspěchu a klíčovým přiřazením |
| 🔐 **Vynucování rozsahu MCP** | 9 podrobných oprávnění pro řízený přístup k nástrojům |
| 📡 **Správa životního cyklu úkolů A2A** | Seznam/filtrování úloh, kontrola událostí/artefaktů, zrušení spuštěných úloh |
| 📋 **Objevení karty agenta** | `/.well-known/agent.json` pro automatické vyhledávání klientů |
| 🧪 **Testovací postroj Protocol E2E** | Skutečné MCP SDK + toky klientů A2A v `test:protocols:e2e` |
| ⚙️ **Provozní kontroly** | Kombinace přepínačů, použití profilů odolnosti, resetování jističů z jednoho ovládacího panelu |
### 🧠 Směrování a inteligence
Funkce | Co to dělá
--- | ---
🎯 **Inteligentní čtyřúrovňový záložní systém** | Automatická trasa: Předplatné → API klíč → Levné → Zdarma
📊 **Sledování kvót v reálném čase** | Počet tokenů v reálném čase + odpočet resetování pro každého poskytovatele
🔄 **Překlad formátu** | OpenAI ↔ Claude ↔ Gemini ↔ Odpovědi s konverzemi bezpečnými pro schéma
👥 **Podpora více účtů** | Více účtů na poskytovatele s inteligentním výběrem
🔄 **Automatická aktualizace tokenů** | Tokeny OAuth se automaticky obnovují při opakovaném pokusu.
🎨 **Vlastní kombinace** | 6 vyvažovacích strategií + řízení záložního řetězce
🌐 **Směrovač se zástupnými znaky** | dynamické směrování `provider/*`
🧠 **Přemýšlení o rozpočtových kontrolách** | Limity pro průchozí, automatické, vlastní a adaptivní uvažování
🔀 **Aliasy modelů** | Vestavěné + vlastní aliasování modelů a bezpečnost migrace
**Degradace pozadí** | Směrujte úlohy na pozadí s nízkou prioritou na levnější modely
🧪 **Chytré směrování s ohledem na úkoly** | Automatický výběr modelu podle typu obsahu (kódování/vize/analýza/sumarizace)
💬 **Vstřikování do systému** | Globální kontroly chování uplatňované konzistentně
📄 **Kompatibilita API pro odpovědi** | Plná podpora `/v1/responses` pro Codex a pokročilé agentické pracovní postupy
| Funkce | Co to dělá |
| ----------------------------------------------- | ----------------------------------------------------------------------------- |
| 🎯 **Inteligentní čtyřúrovňový záložní systém** | Automatická trasa: Předplatné → API klíč → Levné → Zdarma |
| 📊 **Sledování kvót v reálném čase** | Počet tokenů v reálném čase + odpočet resetování pro každého poskytovatele |
| 🔄 **Překlad formátu** | OpenAI ↔ Claude ↔ Gemini ↔ Odpovědi s konverzemi bezpečnými pro schéma |
| 👥 **Podpora více účtů** | Více účtů na poskytovatele s inteligentním výběrem |
| 🔄 **Automatická aktualizace tokenů** | Tokeny OAuth se automaticky obnovují při opakovaném pokusu. |
| 🎨 **Vlastní kombinace** | 6 vyvažovacích strategií + řízení záložního řetězce |
| 🌐 **Směrovač se zástupnými znaky** | dynamické směrování `provider/*` |
| 🧠 **Přemýšlení o rozpočtových kontrolách** | Limity pro průchozí, automatické, vlastní a adaptivní uvažování |
| 🔀 **Aliasy modelů** | Vestavěné + vlastní aliasování modelů a bezpečnost migrace |
| **Degradace pozadí** | Směrujte úlohy na pozadí s nízkou prioritou na levnější modely |
| 🧪 **Chytré směrování s ohledem na úkoly** | Automatický výběr modelu podle typu obsahu (kódování/vize/analýza/sumarizace) |
| 💬 **Vstřikování do systému** | Globální kontroly chování uplatňované konzistentně |
| 📄 **Kompatibilita API pro odpovědi** | Plná podpora `/v1/responses` pro Codex a pokročilé agentické pracovní postupy |
### 🎵 Multimodální API
Funkce | Co to dělá
--- | ---
🖼️ **Generování obrázků** | `/v1/images/generations` s cloudovým a lokálním backendem
📐 **Vložení** | `/v1/embeddings` pro vyhledávání a RAG pipelines
🎤 **Přepis zvuku** | `/v1/audio/transcriptions` (Whisper a další poskytovatelé)
🔊 **Převod textu na řeč** | `/v1/audio/speech` (více enginů/poskytovatelů)
🎬 **Generování videa** | `/v1/videos/generations` (pracovní postupy ComfyUI + SD WebUI)
🎵 **Hudební generace** | `/v1/music/generations` (pracovní postupy ComfyUI)
🛡️ **Moderování** | Bezpečnostní kontroly `/v1/moderations`
🔀 **Změna pořadí** | `/v1/rerank` pro hodnocení relevance
🔍 **Vyhledávání na webu** 🆕 | `/v1/search` — 5 poskytovatelů (Serper, Brave, Perplexity, Exa, Tavily), více než 6 500 zdarma/měsíc, automatické přepnutí na záložní systém, mezipaměť
| Funkce | Co to dělá |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Generování obrázků** | `/v1/images/generations` s cloudovým a lokálním backendem |
| 📐 **Vložení** | `/v1/embeddings` pro vyhledávání a RAG pipelines |
| 🎤 **Přepis zvuku** | `/v1/audio/transcriptions` (Whisper a další poskytovatelé) |
| 🔊 **Převod textu na řeč** | `/v1/audio/speech` (více enginů/poskytovatelů) |
| 🎬 **Generování videa** | `/v1/videos/generations` (pracovní postupy ComfyUI + SD WebUI) |
| 🎵 **Hudební generace** | `/v1/music/generations` (pracovní postupy ComfyUI) |
| 🛡️ **Moderování** | Bezpečnostní kontroly `/v1/moderations` |
| 🔀 **Změna pořadí** | `/v1/rerank` pro hodnocení relevance |
| 🔍 **Vyhledávání na webu** 🆕 | `/v1/search` — 5 poskytovatelů (Serper, Brave, Perplexity, Exa, Tavily), více než 6 500 zdarma/měsíc, automatické přepnutí na záložní systém, mezipaměť |
### 🛡️ Odolnost, bezpečnost a správa věcí veřejných
Funkce | Co to dělá
--- | ---
🔌 **Jističe** | Vypnutí/obnovení pro každý model s ovládáním prahových hodnot
🎯 **Modely s ohledem na koncové body** | Vlastní modely deklarují podporované koncové body + formát API
🛡️ **Stádo proti hromům** | Ochrana mutexu a semaforu při událostech opakování/rychlosti
🧠 **Sémantická + podpisová mezipaměť** | Snížení nákladů/latence díky dvěma vrstvám mezipaměti
**Žádost o idempotenci** | Okno ochrany proti duplikacím
🔒 **Falšování otisků prstů pomocí TLS** | Otisk TLS podobný prohlížeči **snižuje detekci botů a nahlašování účtů**
🔏 **Porovnávání otisků prstů v CLI** | Shoduje se s nativními podpisy požadavků CLI **snižuje riziko zablokování a zároveň zachovává IP adresu proxy**
🌐 **Filtrování IP adres** | Ovládání seznamu povolených/blokovaných položek pro odhalená nasazení
📊 **Upravitelné limity rychlosti** | Konfigurovatelné globální/na úrovni poskytovatele limity s perzistencí
🔑 **Správa klíčů API a stanovení rozsahu** | Bezpečné vydávání/rotace klíčů a kontroly modelu/poskytovatele
🛡️ **Chráněné `/models`** | Volitelné ověřování a skrytí poskytovatele pro katalog modelů
| Funkce | Co to dělá |
| ------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| 🔌 **Jističe** | Vypnutí/obnovení pro každý model s ovládáním prahových hodnot |
| 🎯 **Modely s ohledem na koncové body** | Vlastní modely deklarují podporované koncové body + formát API |
| 🛡️ **Stádo proti hromům** | Ochrana mutexu a semaforu při událostech opakování/rychlosti |
| 🧠 **Sémantická + podpisová mezipaměť** | Snížení nákladů/latence díky dvěma vrstvám mezipaměti |
| **Žádost o idempotenci** | Okno ochrany proti duplikacím |
| 🔒 **Falšování otisků prstů pomocí TLS** | Otisk TLS podobný prohlížeči **snižuje detekci botů a nahlašování účtů** |
| 🔏 **Porovnávání otisků prstů v CLI** | Shoduje se s nativními podpisy požadavků CLI **snižuje riziko zablokování a zároveň zachovává IP adresu proxy** |
| 🌐 **Filtrování IP adres** | Ovládání seznamu povolených/blokovaných položek pro odhalená nasazení |
| 📊 **Upravitelné limity rychlosti** | Konfigurovatelné globální/na úrovni poskytovatele limity s perzistencí |
| 🔑 **Správa klíčů API a stanovení rozsahu** | Bezpečné vydávání/rotace klíčů a kontroly modelu/poskytovatele |
| 🛡️ **Chráněné `/models`** | Volitelné ověřování a skrytí poskytovatele pro katalog modelů |
### 📊 Pozorovatelnost a analytika
Funkce | Co to dělá
--- | ---
📝 **Žádost + protokolování proxy** | Úplné protokolování požadavků/odpovědí a proxy
📋 **Sjednocený panel protokolů** | Zobrazení požadavků, proxy, auditu a konzole na jedné stránce
🔍 **Vyžádat si telemetrii** | Latence p50/p95/p99 a trasování požadavků
🏥 **Panel zdraví** | Doba provozuschopnosti, stavy jističů, uzamčení, statistiky mezipaměti
💰 **Sledování nákladů** | Kontrola rozpočtu a přehled o cenách pro jednotlivé modely
📈 **Analytické vizualizace** | Přehledy využití modelů/poskytovatelů a zobrazení trendů
🧪 **Rámec hodnocení** | Testování zlaté sady s konfigurovatelnými strategiemi shody
| Funkce | Co to dělá |
| ----------------------------------- | ---------------------------------------------------------------------- |
| 📝 **Žádost + protokolování proxy** | Úplné protokolování požadavků/odpovědí a proxy |
| 📋 **Sjednocený panel protokolů** | Zobrazení požadavků, proxy, auditu a konzole na jedné stránce |
| 🔍 **Vyžádat si telemetrii** | Latence p50/p95/p99 a trasování požadavků |
| 🏥 **Panel zdraví** | Doba provozuschopnosti, stavy jističů, uzamčení, statistiky mezipaměti |
| 💰 **Sledování nákladů** | Kontrola rozpočtu a přehled o cenách pro jednotlivé modely |
| 📈 **Analytické vizualizace** | Přehledy využití modelů/poskytovatelů a zobrazení trendů |
| 🧪 **Rámec hodnocení** | Testování zlaté sady s konfigurovatelnými strategiemi shody |
### ☁️ Nasazení a platforma
Funkce | Co to dělá
--- | ---
🌐 **Nasazení kdekoli** | Localhost, VPS, Docker, cloudová prostředí
💾 **Synchronizace s cloudem** | Synchronizace konfigurace přes cloud worker
🔄 **Zálohování/Obnovení** | Toky exportu/importu a obnovy po havárii
🧙 **Průvodce nástupem** | Průvodce prvním spuštěním
🔧 **Panel nástrojů CLI** | Nastavení oblíbených kódovacích nástrojů jedním kliknutím
🎮 **Modelové hřiště** | Otestujte libovolného poskytovatele/model/koncový bod z řídicího panelu
🔏 **Přepínač otisků prstů v příkazovém řádku** | Porovnávání otisků prstů podle poskytovatele v Nastavení &gt; Zabezpečení
🌐 **i18n (30 jazyků)** | Plná jazyková podpora dashboardu a dokumentace s psaním zprava doleva
📂 **Adresář vlastních dat** | Přepsání `DATA_DIR` pro umístění úložiště
| Funkce | Co to dělá |
| ----------------------------------------------- | ------------------------------------------------------------------------- |
| 🌐 **Nasazení kdekoli** | Localhost, VPS, Docker, cloudová prostředí |
| 💾 **Synchronizace s cloudem** | Synchronizace konfigurace přes cloud worker |
| 🔄 **Zálohování/Obnovení** | Toky exportu/importu a obnovy po havárii |
| 🧙 **Průvodce nástupem** | Průvodce prvním spuštěním |
| 🔧 **Panel nástrojů CLI** | Nastavení oblíbených kódovacích nástrojů jedním kliknutím |
| 🎮 **Modelové hřiště** | Otestujte libovolného poskytovatele/model/koncový bod z řídicího panelu |
| 🔏 **Přepínač otisků prstů v příkazovém řádku** | Porovnávání otisků prstů podle poskytovatele v Nastavení &gt; Zabezpečení |
| 🌐 **i18n (30 jazyků)** | Plná jazyková podpora dashboardu a dokumentace s psaním zprava doleva |
| 📂 **Adresář vlastních dat** | Přepsání `DATA_DIR` pro umístění úložiště |
### Hluboký pohled na funkce
@@ -1203,12 +1115,12 @@ Předinstalovaná sada „OmniRoute Golden Set“ obsahuje testovací případy
### Strategie hodnocení
Strategie | Popis | Příklad
--- | --- | ---
`exact` | Výstup se musí přesně shodovat | `"4"`
`contains` | Výstup musí obsahovat podřetězec (bez rozlišení velkých a malých písmen) | `"Paris"`
`regex` | Výstup musí odpovídat vzoru regulárních výrazů | `"1.*2.*3"`
`custom` | Vlastní JS funkce vrací true/false | `(output) => output.length > 10`
| Strategie | Popis | Příklad |
| ---------- | ------------------------------------------------------------------------ | -------------------------------- |
| `exact` | Výstup se musí přesně shodovat | `"4"` |
| `contains` | Výstup musí obsahovat podřetězec (bez rozlišení velkých a malých písmen) | `"Paris"` |
| `regex` | Výstup musí odpovídat vzoru regulárních výrazů | `"1.*2.*3"` |
| `custom` | Vlastní JS funkce vrací true/false | `(output) => output.length > 10` |
---
@@ -1240,9 +1152,6 @@ Užitečná API pro automatizaci:
- `GET /api/mcp/audit`
- `GET /api/mcp/audit/stats`
<details>
<summary><b>🤝 Nastavení A2A (Agent2Agent)</b></summary>
</details>
@@ -1272,9 +1181,6 @@ Provozní uživatelské rozhraní:
- `/dashboard/a2a` pro pozorovatelnost úloh/stavů/streamů a akce kouření
<details>
<summary><b>🧪 Komplexní validace protokolu</b></summary>
</details>
@@ -1291,9 +1197,6 @@ Tím se ověřuje:
- A2A objevování/odesílání/streamování/získávání/zrušení
- Křížová kontrola dat v auditu MCP a API pro správu úloh A2A
<details>
<summary><b>💳 Poskytovatelé předplatného</b></summary>
</details>
@@ -1369,9 +1272,6 @@ Models:
gh/gemini-3-pro
```
<details>
<summary><b>🔑 Poskytovatelé klíčů API</b></summary>
</details>
@@ -1381,7 +1281,7 @@ Models:
1. Registrace: [build.nvidia.com](https://build.nvidia.com)
2. Získejte zdarma klíč API (včetně 1000 inferenčních kreditů)
3. Ovládací panel → Přidat poskytovatele → NVIDIA NIM:
- Klíč API: `nvapi-your-key`
- Klíč API: `nvapi-your-key`
**Modely:** `nvidia/llama-3.3-70b-instruct` , `nvidia/mistral-7b-instruct` a více než 50 dalších
@@ -1413,9 +1313,6 @@ Models:
**Modely:** Získejte přístup k více než 100 modelům od všech hlavních poskytovatelů prostřednictvím jediného klíče API.
<details>
<summary><b>💰 Levní poskytovatelé (záložní)</b></summary>
</details>
@@ -1425,8 +1322,8 @@ Models:
1. Registrace: [Zhipu AI](https://open.bigmodel.cn/)
2. Získejte klíč API z kódovacího plánu
3. Nástěnka → Přidat klíč API:
- Poskytovatel: `glm`
- Klíč API: `your-key`
- Poskytovatel: `glm`
- Klíč API: `your-key`
**Použití:** `glm/glm-4.7`
@@ -1452,9 +1349,6 @@ Models:
**Tip pro profesionály:** Fixních 9 $/měsíc za 10 milionů tokenů = efektivní náklady 0,90 $/1 milion!
<details>
<summary><b>🆓 BEZPLATNÍ poskytovatelé (nouzové zálohování)</b></summary>
</details>
@@ -1498,9 +1392,6 @@ Models:
kr/claude-haiku-4.5
```
<details>
<summary><b>🎨 Vytvořte kombinace</b></summary>
</details>
@@ -1531,9 +1422,6 @@ Models:
Cost: $0 forever!
```
<details>
<summary><b>🔧 Integrace s rozhraním příkazového řádku</b></summary>
</details>
@@ -1637,9 +1525,6 @@ opencode
> **Tip:** Do sekce `models` přidejte jakýkoli model dostupný ve vašem koncovém bodu OmniRoute `/v1/models` . Použijte formát `provider/model-id` z vašeho dashboardu OmniRoute.
---
## 🐛 Řešení problémů
@@ -1880,14 +1765,8 @@ Chcete-li získat přístup k kriterii pověření, můžete použít adresu **U
> Toto řešení funguje na základě autorizačního kódu na adrese URL a nezávislého přesměrování přesměrování nebo jiného.
---
## 🛠️ Technologický stack
<details>
@@ -1909,28 +1788,25 @@ Chcete-li získat přístup k kriterii pověření, můžete použít adresu **U
- **Docker** : [hub.docker.com/r/diegosouzapw/omniroute](https://hub.docker.com/r/diegosouzapw/omniroute)
- **Odolnost** : Jistič, exponenciální odstavení, ochrana proti hromům, falešné TLS, automatické kombinované samoopravování
---
## 📖 Dokumentace
Dokument | Popis
--- | ---
[Uživatelská příručka](docs/USER_GUIDE.md) | Poskytovatelé, kombinace, integrace CLI, nasazení
[Referenční informace k API](docs/API_REFERENCE.md) | Všechny koncové body s příklady
[MCP server](open-sse/mcp-server/README.md) | 16 nástrojů MCP, konfigurace IDE, klienti Python/TS/Go
[Server A2A](src/lib/a2a/README.md) | Protokol JSON-RPC 2.0, dovednosti, streamování, správa úloh
[Auto-Combo Engine](docs/auto-combo.md) | 6faktorové bodování, balíčky režimů, samoléčba
[Odstraňování problémů](docs/TROUBLESHOOTING.md) | Běžné problémy a jejich řešení
[Architektura](docs/ARCHITECTURE.md) | Architektura a interní prvky systému
[Přispívání](CONTRIBUTING.md) | Nastavení a pokyny pro vývoj
[Specifikace OpenAPI](docs/openapi.yaml) | Specifikace OpenAPI 3.0
[Bezpečnostní zásady](SECURITY.md) | Hlášení zranitelností a bezpečnostní postupy
[Nasazení virtuálního počítače](docs/VM_DEPLOYMENT_GUIDE.md) | Kompletní průvodce: Nastavení virtuálního počítače + nginx + Cloudflare
[Galerie funkcí](docs/FEATURES.md) | Vizuální prohlídka řídicího panelu se snímky obrazovky
[Kontrolní seznam vydání](docs/RELEASE_CHECKLIST.md) | Kroky ověření před vydáním
| Dokument | Popis |
| ------------------------------------------------------------ | ----------------------------------------------------------------------- |
| [Uživatelská příručka](docs/USER_GUIDE.md) | Poskytovatelé, kombinace, integrace CLI, nasazení |
| [Referenční informace k API](docs/API_REFERENCE.md) | Všechny koncové body s příklady |
| [MCP server](open-sse/mcp-server/README.md) | 16 nástrojů MCP, konfigurace IDE, klienti Python/TS/Go |
| [Server A2A](src/lib/a2a/README.md) | Protokol JSON-RPC 2.0, dovednosti, streamování, správa úloh |
| [Auto-Combo Engine](docs/auto-combo.md) | 6faktorové bodování, balíčky režimů, samoléčba |
| [Odstraňování problémů](docs/TROUBLESHOOTING.md) | Běžné problémy a jejich řešení |
| [Architektura](docs/ARCHITECTURE.md) | Architektura a interní prvky systému |
| [Přispívání](CONTRIBUTING.md) | Nastavení a pokyny pro vývoj |
| [Specifikace OpenAPI](docs/openapi.yaml) | Specifikace OpenAPI 3.0 |
| [Bezpečnostní zásady](SECURITY.md) | Hlášení zranitelností a bezpečnostní postupy |
| [Nasazení virtuálního počítače](docs/VM_DEPLOYMENT_GUIDE.md) | Kompletní průvodce: Nastavení virtuálního počítače + nginx + Cloudflare |
| [Galerie funkcí](docs/FEATURES.md) | Vizuální prohlídka řídicího panelu se snímky obrazovky |
| [Kontrolní seznam vydání](docs/RELEASE_CHECKLIST.md) | Kroky ověření před vydáním |
---
@@ -1938,14 +1814,14 @@ Dokument | Popis
OmniRoute má **v plánu více než 210 funkcí** v několika fázích vývoje. Zde jsou klíčové oblasti:
Kategorie | Plánované funkce | Hlavní body
--- | --- | ---
🧠 **Směrování a inteligence** | 25+ | Směrování s nejnižší latencí, směrování založené na tagech, kontrola kvót před výstupem, výběr účtu P2C
🔒 **Zabezpečení a dodržování předpisů** | 20+ | Zpevnění SSRF, maskování přihlašovacích údajů, limit rychlosti pro každý koncový bod, stanovení rozsahu klíčů pro správu
📊 **Pozorovatelnost** | 15+ | Integrace OpenTelemetry, sledování kvót v reálném čase, sledování nákladů podle modelu
🔄 **Integrace poskytovatelů** | 20+ | Dynamický registr modelů, doba zchlazení poskytovatelů, Codex pro více účtů, analýza kvót Copilota
**Výkon** | 15+ | Dvojitá vrstva mezipaměti, mezipaměť výzev, mezipaměť odpovědí, udržování streamování, dávkové API
🌐 **Ekosystém** | 10+ | WebSocket API, horké opětovné načítání konfigurace, distribuované úložiště konfigurace, komerční režim
| Kategorie | Plánované funkce | Hlavní body |
| ---------------------------------------- | ---------------- | ------------------------------------------------------------------------------------------------------------------------ |
| 🧠 **Směrování a inteligence** | 25+ | Směrování s nejnižší latencí, směrování založené na tagech, kontrola kvót před výstupem, výběr účtu P2C |
| 🔒 **Zabezpečení a dodržování předpisů** | 20+ | Zpevnění SSRF, maskování přihlašovacích údajů, limit rychlosti pro každý koncový bod, stanovení rozsahu klíčů pro správu |
| 📊 **Pozorovatelnost** | 15+ | Integrace OpenTelemetry, sledování kvót v reálném čase, sledování nákladů podle modelu |
| 🔄 **Integrace poskytovatelů** | 20+ | Dynamický registr modelů, doba zchlazení poskytovatelů, Codex pro více účtů, analýza kvót Copilota |
| **Výkon** | 15+ | Dvojitá vrstva mezipaměti, mezipaměť výzev, mezipaměť odpovědí, udržování streamování, dávkové API |
| 🌐 **Ekosystém** | 10+ | WebSocket API, horké opětovné načítání konfigurace, distribuované úložiště konfigurace, komerční režim |
### 🔜 Již brzy
+2 -2
View File
@@ -934,8 +934,8 @@ OmniRoute v2.0 er bygget som en operationel platform, ikke kun en relæ-proxy.
| Funktion | Hvad det gør || -------------------------- | -------------------------------------------------------------------- |
| 🖼️ **Billedgenerering** | `/v1/images/generations` med cloud og lokale backends |
| 📐 **Indlejringer** | `/v1/embeddings` til søgning og RAG-rørledninger |
| 🎤 **Lydtransskription** | `/v1/audio/transcriptions` (Whisper og yderligere udbydere) |
| 🔊 **Tekst-til-tale** | `/v1/audio/speech` (flere motorer/udbydere) |
| 🎤 **Lydtransskription** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Tekst-til-tale** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🎬 **Videogenerering** | `/v1/videos/generations` (ComfyUI + SD WebUI-arbejdsgange) |
| 🎵 **Music Generation** | `/v1/music/generations` (ComfyUI-arbejdsgange) |
| 🛡️ **Moderationer** | `/v1/moderations` sikkerhedstjek |
+2 -2
View File
@@ -939,8 +939,8 @@ OmniRoute v2.0 ist als Betriebsplattform konzipiert und nicht nur als Relay-Prox
| Funktion | Was es tut || -------------------------- | ------------------------------------------------------------- |
| 🖼️ **Bilderzeugung** | `/v1/images/generations` mit Cloud- und lokalen Backends |
| 📐 **Einbettungen** | `/v1/embeddings` für Such- und RAG-Pipelines |
| 🎤 **Audio-Transkription** | `/v1/audio/transcriptions` (Whisper und zusätzliche Anbieter) |
| 🔊 **Text-to-Speech** | `/v1/audio/speech` (mehrere Engines/Anbieter) |
| 🎤 **Audio-Transkription** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Text-to-Speech** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🎬 **Videogenerierung** | `/v1/videos/generations` (ComfyUI + SD WebUI-Workflows) |
| 🎵 **Musikgeneration** | `/v1/music/generations` (ComfyUI-Workflows) |
| 🛡️ **Moderationen** | `/v1/moderations` Sicherheitsprüfungen |
+8 -8
View File
@@ -877,14 +877,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 APIs Multi-Modal
| Característica | Qué Hace |
| ----------------------------- | ------------------------------------------------------ |
| 🖼️ **Generación de Imágenes** | `/v1/images/generations` — 4 proveedores, 9+ modelos |
| 📐 **Embeddings** | `/v1/embeddings` — 6 proveedores, 9+ modelos |
| 🎤 **Transcripción de Audio** | `/v1/audio/transcriptions`Compatible con Whisper |
| 🔊 **Texto a Voz** | `/v1/audio/speech`Síntesis de audio multi-proveedor |
| 🛡️ **Moderaciones** | `/v1/moderations` — Verificaciones de seguridad |
| 🔀 **Reranking** | `/v1/rerank` — Reranking de relevancia de documentos |
| Característica | Qué Hace |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Generación de Imágenes** | `/v1/images/generations` — 4 proveedores, 9+ modelos |
| 📐 **Embeddings** | `/v1/embeddings` — 6 proveedores, 9+ modelos |
| 🎤 **Transcripción de Audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Texto a Voz** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderaciones** | `/v1/moderations` — Verificaciones de seguridad |
| 🔀 **Reranking** | `/v1/rerank` — Reranking de relevancia de documentos |
### 🛡️ Resiliencia y Seguridad
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multimodaaliset sovellusliittymät
| Ominaisuus | Mitä se tekee |
| ------------------------- | --------------------------------------------------------- |
| 🖼️ **Kuvan luominen** | `/v1/images/generations` — 4 toimittajaa, 9+ mallia |
| 📐 **Upotukset** | `/v1/embeddings` — 6 toimittajaa, 9+ mallia |
| 🎤 **Äänitranskriptio** | `/v1/audio/transcriptions`Kuiskausyhteensopiva |
| 🔊 **Tekstistä puheeksi** | `/v1/audio/speech`Usean palveluntarjoajan äänisynteesi |
| 🛡️ **Moderaatiot** | `/v1/moderations` — Sisällön turvallisuustarkistukset |
| 🔀 **Uudelleenjärjestys** | `/v1/rerank` — Asiakirjan osuvuuden uudelleensijoitus |
| Ominaisuus | Mitä se tekee |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Kuvan luominen** | `/v1/images/generations` — 4 toimittajaa, 9+ mallia |
| 📐 **Upotukset** | `/v1/embeddings` — 6 toimittajaa, 9+ mallia |
| 🎤 **Äänitranskriptio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Tekstistä puheeksi** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderaatiot** | `/v1/moderations` — Sisällön turvallisuustarkistukset |
| 🔀 **Uudelleenjärjestys** | `/v1/rerank` — Asiakirjan osuvuuden uudelleensijoitus |
### 🛡️ Joustavuus ja turvallisuus
+8 -8
View File
@@ -875,14 +875,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 APIs multi-modales
| Fonctionnalité | Ce qu'elle fait |
| -------------------------- | ------------------------------------------------------- |
| 🖼️ **Génération d'images** | `/v1/images/generations` — 4 fournisseurs, 9+ modèles |
| 📐 **Embeddings** | `/v1/embeddings` — 6 fournisseurs, 9+ modèles |
| 🎤 **Transcription audio** | `/v1/audio/transcriptions`compatible Whisper |
| 🔊 **Texte vers parole** | `/v1/audio/speech`synthèse audio multi-fournisseur |
| 🛡️ **Modérations** | `/v1/moderations` — vérifications de sécurité |
| 🔀 **Reranking** | `/v1/rerank` — reclassement de pertinence des documents |
| Fonctionnalité | Ce qu'elle fait |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Génération d'images** | `/v1/images/generations` — 4 fournisseurs, 9+ modèles |
| 📐 **Embeddings** | `/v1/embeddings` — 6 fournisseurs, 9+ modèles |
| 🎤 **Transcription audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Texte vers parole** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Modérations** | `/v1/moderations` — vérifications de sécurité |
| 🔀 **Reranking** | `/v1/rerank` — reclassement de pertinence des documents |
### 🛡️ Résilience & Sécurité
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 ממשקי API רב-מודאליים
| תכונה | מה זה עושה |
| ------------------- | --------------------------------------------- |
| 🖼️ **יצירת תמונות** | `/v1/images/generations` — 4 ספקים, 9+ דגמים |
| 📐 **הטבעות** | `/v1/embeddings` — 6 ספקים, 9+ דגמים |
| 🎤 **תמלול אודיו** | `/v1/audio/transcriptions`תואם לחישה |
| 🔊 **טקסט לדיבור** | `/v1/audio/speech`סינתזת אודיו מרובה ספקים |
| 🛡️ **מנחים** | `/v1/moderations` — בדיקות בטיחות תוכן |
| 🔀 **דירוג מחדש** | `/v1/rerank` — דירוג מחדש של רלוונטיות המסמך |
| תכונה | מה זה עושה |
| ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **יצירת תמונות** | `/v1/images/generations` — 4 ספקים, 9+ דגמים |
| 📐 **הטבעות** | `/v1/embeddings` — 6 ספקים, 9+ דגמים |
| 🎤 **תמלול אודיו** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **טקסט לדיבור** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **מנחים** | `/v1/moderations` — בדיקות בטיחות תוכן |
| 🔀 **דירוג מחדש** | `/v1/rerank` — דירוג מחדש של רלוונטיות המסמך |
### 🛡️ חוסן וביטחון
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multimodális API-k
| Funkció | Mit csinál |
| ---------------------- | ------------------------------------------------------------ |
| 🖼️ **Képgenerálás** | `/v1/images/generations` — 4 szolgáltató, 9+ modell |
| 📐 **Beágyazás** | `/v1/embeddings` — 6 szolgáltató, 9+ modell |
| 🎤 **Audio átírás** | `/v1/audio/transcriptions`Suttogás-kompatibilis |
| 🔊 **Szövegfelolvasó** | `/v1/audio/speech`Hangszintézis több szolgáltatónál |
| 🛡️ **Moderálás** | `/v1/moderations` — Tartalombiztonsági ellenőrzések |
| 🔀 **Átsorolás** | `/v1/rerank` — A dokumentumok relevancia szerinti átsorolása |
| Funkció | Mit csinál |
| ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Képgenerálás** | `/v1/images/generations` — 4 szolgáltató, 9+ modell |
| 📐 **Beágyazás** | `/v1/embeddings` — 6 szolgáltató, 9+ modell |
| 🎤 **Audio átírás** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Szövegfelolvasó** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderálás** | `/v1/moderations` — Tartalombiztonsági ellenőrzések |
| 🔀 **Átsorolás** | `/v1/rerank` — A dokumentumok relevancia szerinti átsorolása |
### 🛡️ Rugalmasság és biztonság
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 API Multi-Modal
| Fitur | Apa Fungsinya |
| -------------------------- | ------------------------------------------------------ |
| 🖼️ **Pembuatan Gambar** | `/v1/images/generations` — 4 penyedia, 9+ model |
| 📐 **Sematan** | `/v1/embeddings` — 6 penyedia, 9+ model |
| 🎤 **Transkripsi Audio** | `/v1/audio/transcriptions`Kompatibel dengan bisikan |
| 🔊 **Teks-ke-Ucapan** | `/v1/audio/speech`Sintesis audio multi-penyedia |
| 🛡️ **Moderasi** | `/v1/moderations` — Pemeriksaan keamanan konten |
| 🔀 **Pemeringkatan Ulang** | `/v1/rerank` — Pemeringkatan ulang relevansi dokumen |
| Fitur | Apa Fungsinya |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Pembuatan Gambar** | `/v1/images/generations` — 4 penyedia, 9+ model |
| 📐 **Sematan** | `/v1/embeddings` — 6 penyedia, 9+ model |
| 🎤 **Transkripsi Audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Teks-ke-Ucapan** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderasi** | `/v1/moderations` — Pemeriksaan keamanan konten |
| 🔀 **Pemeringkatan Ulang** | `/v1/rerank` — Pemeringkatan ulang relevansi dokumen |
### 🛡️ Ketahanan & Keamanan
+8 -8
View File
@@ -770,14 +770,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 मल्टी-मॉडल एपीआई
| फ़ीचर | यह क्या करता है |
| ---------------------------- | ------------------------------------------------- |
| 🖼️ **छवि निर्माण** | `/v1/images/generations` - 4 प्रदाता, 9+ मॉडल |
| 📐 **एंबेडिंग** | `/v1/embeddings` — 6 प्रदाता, 9+ मॉडल |
| 🎤 **ऑडियो ट्रांस्क्रिप्शन** | `/v1/audio/transcriptions` - कानाफूसी-संगत |
| 🔊 **टेक्स्ट-टू-स्पीच** | `/v1/audio/speech` - बहु-प्रदाता ऑडियो संश्लेषण |
| 🛡️ **संयम** | `/v1/moderations` — सामग्री सुरक्षा जांच |
| 🔀 **पुनर्रैंकिंग** | `/v1/rerank` — दस्तावेज़ प्रासंगिकता पुनर्रैंकिंग |
| फ़ीचर | यह क्या करता है |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **छवि निर्माण** | `/v1/images/generations` - 4 प्रदाता, 9+ मॉडल |
| 📐 **एंबेडिंग** | `/v1/embeddings` — 6 प्रदाता, 9+ मॉडल |
| 🎤 **ऑडियो ट्रांस्क्रिप्शन** | `/v1/audio/transcriptions` — 7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **टेक्स्ट-टू-स्पीच** | `/v1/audio/speech` — 10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **संयम** | `/v1/moderations` — सामग्री सुरक्षा जांच |
| 🔀 **पुनर्रैंकिंग** | `/v1/rerank` — दस्तावेज़ प्रासंगिकता पुनर्रैंकिंग |
### 🛡️ लचीलापन और सुरक्षा
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 API Multi-modali
| Funzionalità | Cosa Fa |
| --------------------------- | ---------------------------------------------------- |
| 🖼️ **Generazione immagini** | `/v1/images/generations` — 4 provider, 9+ modelli |
| 📐 **Embeddings** | `/v1/embeddings` — 6 provider, 9+ modelli |
| 🎤 **Trascrizione audio** | `/v1/audio/transcriptions`Compatibile Whisper |
| 🔊 **Testo a voce** | `/v1/audio/speech`Sintesi audio multi-provider |
| 🛡️ **Moderazioni** | `/v1/moderations` — Controlli di sicurezza |
| 🔀 **Reranking** | `/v1/rerank` — Riclassificazione rilevanza documenti |
| Funzionalità | Cosa Fa |
| --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Generazione immagini** | `/v1/images/generations` — 4 provider, 9+ modelli |
| 📐 **Embeddings** | `/v1/embeddings` — 6 provider, 9+ modelli |
| 🎤 **Trascrizione audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Testo a voce** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderazioni** | `/v1/moderations` — Controlli di sicurezza |
| 🔀 **Reranking** | `/v1/rerank` — Riclassificazione rilevanza documenti |
### 🛡️ Resilienza & Sicurezza
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 マルチモーダル API
| 特集 | 何をするのか |
| ----------------------- | --------------------------------------------------------------- |
| 🖼️ **画像生成** | `/v1/images/generations` — 4 つのプロバイダー、9 つ以上のモデル |
| 📐 **埋め込み** | `/v1/embeddings` — 6 つのプロバイダー、9 つ以上のモデル |
| 🎤 **音声文字起こし** | `/v1/audio/transcriptions`ウィスパー互換 |
| 🔊 **テキスト読み上げ** | `/v1/audio/speech`マルチプロバイダーのオーディオ合成 |
| 🛡️ **モデレーション** | `/v1/moderations` — コンテンツの安全性チェック |
| 🔀 **再ランキング** | `/v1/rerank` — ドキュメントの関連性の再ランキング |
| 特集 | 何をするのか |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **画像生成** | `/v1/images/generations` — 4 つのプロバイダー、9 つ以上のモデル |
| 📐 **埋め込み** | `/v1/embeddings` — 6 つのプロバイダー、9 つ以上のモデル |
| 🎤 **音声文字起こし** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **テキスト読み上げ** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **モデレーション** | `/v1/moderations` — コンテンツの安全性チェック |
| 🔀 **再ランキング** | `/v1/rerank` — ドキュメントの関連性の再ランキング |
### 🛡️ 復元力とセキュリティ
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 다중 모드 API
| 기능 | 그것이 하는 일 |
| ----------------------- | ------------------------------------------------------ |
| 🖼️ **이미지 생성** | `/v1/images/generations` — 4개 공급자, 9개 이상의 모델 |
| 📐 **임베딩** | `/v1/embeddings` — 6개 공급자, 9개 이상의 모델 |
| 🎤 **오디오 전사** | `/v1/audio/transcriptions`속삭임 호환 |
| 🔊 **텍스트 음성 변환** | `/v1/audio/speech`다중 제공자 오디오 합성 |
| 🛡️ **조정** | `/v1/moderations` — 콘텐츠 안전 확인 |
| 🔀 **재순위** | `/v1/rerank` — 문서 관련성 재순위 |
| 기능 | 그것이 하는 일 |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **이미지 생성** | `/v1/images/generations` — 4개 공급자, 9개 이상의 모델 |
| 📐 **임베딩** | `/v1/embeddings` — 6개 공급자, 9개 이상의 모델 |
| 🎤 **오디오 전사** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **텍스트 음성 변환** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **조정** | `/v1/moderations` — 콘텐츠 안전 확인 |
| 🔀 **재순위** | `/v1/rerank` — 문서 관련성 재순위 |
### 🛡️ 복원력 및 보안
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 API Berbilang Modal
| Ciri | Apa yang Dilakukan |
| ------------------------ | ------------------------------------------------------ |
| 🖼️ **Penjanaan Imej** | `/v1/images/generations` — 4 pembekal, 9+ model |
| 📐 **Pembenaman** | `/v1/embeddings` — 6 pembekal, 9+ model |
| 🎤 **Transkripsi Audio** | `/v1/audio/transcriptions`Serasi dengan bisikan |
| 🔊 **Teks-ke-Ucapan** | `/v1/audio/speech`Sintesis audio berbilang pembekal |
| 🛡️ **Kesederhanaan** | `/v1/moderations` — Pemeriksaan keselamatan kandungan |
| 🔀 **Penyusunan semula** | `/v1/rerank` — Penarafan semula perkaitan dokumen |
| Ciri | Apa yang Dilakukan |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Penjanaan Imej** | `/v1/images/generations` — 4 pembekal, 9+ model |
| 📐 **Pembenaman** | `/v1/embeddings` — 6 pembekal, 9+ model |
| 🎤 **Transkripsi Audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Teks-ke-Ucapan** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Kesederhanaan** | `/v1/moderations` — Pemeriksaan keselamatan kandungan |
| 🔀 **Penyusunan semula** | `/v1/rerank` — Penarafan semula perkaitan dokumen |
### 🛡️ Ketahanan & Keselamatan
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multimodale API's
| Kenmerk | Wat het doet |
| ------------------------ | --------------------------------------------------------- |
| 🖼️ **Beeldgeneratie** | `/v1/images/generations` — 4 providers, 9+ modellen |
| 📐 **Insluitingen** | `/v1/embeddings` — 6 providers, 9+ modellen |
| 🎤 **Audiotranscriptie** | `/v1/audio/transcriptions`Whisper-compatibel |
| 🔊 **Tekst-naar-spraak** | `/v1/audio/speech`Audiosynthese van meerdere providers |
| 🛡️ **Moderaties** | `/v1/moderations` — Veiligheidscontroles van inhoud |
| 🔀 **Herschikking** | `/v1/rerank` — Herschikking van documentrelevantie |
| Kenmerk | Wat het doet |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Beeldgeneratie** | `/v1/images/generations` — 4 providers, 9+ modellen |
| 📐 **Insluitingen** | `/v1/embeddings` — 6 providers, 9+ modellen |
| 🎤 **Audiotranscriptie** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Tekst-naar-spraak** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderaties** | `/v1/moderations` — Veiligheidscontroles van inhoud |
| 🔀 **Herschikking** | `/v1/rerank` — Herschikking van documentrelevantie |
### 🛡️ Veerkracht en veiligheid
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multi-Modal APIer
| Funksjon | Hva det gjør |
| ----------------------- | ------------------------------------------------------ |
| 🖼️ **Bildegenerering** | `/v1/images/generations` — 4 leverandører, 9+ modeller |
| 📐 **Innbygging** | `/v1/embeddings` — 6 leverandører, 9+ modeller |
| 🎤 **Lydtranskripsjon** | `/v1/audio/transcriptions`Whisper-kompatibel |
| 🔊 **Tekst-til-tale** | `/v1/audio/speech`Multi-leverandør lydsyntese |
| 🛡️ **Moderasjoner** | `/v1/moderations` — Innholdssikkerhetssjekker |
| 🔀 **Omrangering** | `/v1/rerank` — Rerangering av dokumentrelevans |
| Funksjon | Hva det gjør |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Bildegenerering** | `/v1/images/generations` — 4 leverandører, 9+ modeller |
| 📐 **Innbygging** | `/v1/embeddings` — 6 leverandører, 9+ modeller |
| 🎤 **Lydtranskripsjon** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Tekst-til-tale** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderasjoner** | `/v1/moderations` — Innholdssikkerhetssjekker |
| 🔀 **Omrangering** | `/v1/rerank` — Rerangering av dokumentrelevans |
### 🛡️ Spenst og sikkerhet
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Mga Multi-Modal na API
| Tampok | Ano ang Ginagawa Nito |
| -------------------------- | ------------------------------------------------------------ |
| 🖼️ **Pagbuo ng Larawan** | `/v1/images/generations` — 4 na provider, 9+ na modelo |
| 📐 **Mga Pag-embed** | `/v1/embeddings` — 6 na provider, 9+ na modelo |
| 🎤 **Audio Transcription** | `/v1/audio/transcriptions`Whisper-compatible |
| 🔊 **Text-to-Speech** | `/v1/audio/speech`Multi-provider audio synthesis |
| 🛡️ **Mga Pag-moderate** | `/v1/moderations` — Mga pagsusuri sa kaligtasan ng nilalaman |
| 🔀 **Reranking** | `/v1/rerank` — Muling pagraranggo ng kaugnayan ng dokumento |
| Tampok | Ano ang Ginagawa Nito |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Pagbuo ng Larawan** | `/v1/images/generations` — 4 na provider, 9+ na modelo |
| 📐 **Mga Pag-embed** | `/v1/embeddings` — 6 na provider, 9+ na modelo |
| 🎤 **Audio Transcription** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Text-to-Speech** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Mga Pag-moderate** | `/v1/moderations` — Mga pagsusuri sa kaligtasan ng nilalaman |
| 🔀 **Reranking** | `/v1/rerank` — Muling pagraranggo ng kaugnayan ng dokumento |
### 🛡️ Katatagan at Seguridad
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Wielomodalne interfejsy API
| Funkcja | Co to robi |
| ----------------------------- | ------------------------------------------------------ |
| 🖼️ **Generowanie obrazu** | `/v1/images/generations` — 4 dostawców, ponad 9 modeli |
| 📐 **Osadzenia** | `/v1/embeddings` — 6 dostawców, ponad 9 modeli |
| 🎤 **Transkrypcja audio** | `/v1/audio/transcriptions`Kompatybilny z szeptem |
| 🔊 **Zamiana tekstu na mowę** | `/v1/audio/speech`Synteza dźwięku wielu dostawców |
| 🛡️ **Moderacje** | `/v1/moderations` — Kontrola bezpieczeństwa treści |
| 🔀 **Ponowna pozycja** | `/v1/rerank` — Zmiana rankingu trafności dokumentu |
| Funkcja | Co to robi |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Generowanie obrazu** | `/v1/images/generations` — 4 dostawców, ponad 9 modeli |
| 📐 **Osadzenia** | `/v1/embeddings` — 6 dostawców, ponad 9 modeli |
| 🎤 **Transkrypcja audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Zamiana tekstu na mowę** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderacje** | `/v1/moderations` — Kontrola bezpieczeństwa treści |
| 🔀 **Ponowna pozycja** | `/v1/rerank` — Zmiana rankingu trafności dokumentu |
### 🛡️ Odporność i bezpieczeństwo
+79 -28
View File
@@ -819,24 +819,28 @@ Quando minimizado, o OmniRoute fica na bandeja do sistema com ações rápidas:
## 💰 Preços Resumidos
| Tier | Provedor | Custo | Reset de Cota | Melhor Para |
| ----------------- | ----------------- | ---------------------------- | ----------------- | ----------------------- |
| **💳 ASSINATURA** | Claude Code (Pro) | $20/mês | 5h + semanal | Já é assinante |
| | Codex (Plus/Pro) | $20-200/mês | 5h + semanal | Usuários OpenAI |
| | Gemini CLI | **GRATUITO** | 180K/mês + 1K/dia | Todos! |
| | GitHub Copilot | $10-19/mês | Mensal | Usuários GitHub |
| **🔑 API KEY** | NVIDIA NIM | **GRATUITO** (1000 créditos) | Único | Testes gratuitos |
| | DeepSeek | Por uso | Nenhum | Melhor preço/qualidade |
| | Groq | Tier gratuito + pago | Limitado | Inferência ultra-rápida |
| | xAI (Grok) | Por uso | Nenhum | Modelos Grok |
| | Mistral | Tier gratuito + pago | Limitado | IA Europeia |
| | OpenRouter | Por uso | Nenhum | 100+ modelos |
| **💰 BARATO** | GLM-4.7 | $0.6/1M | Diário 10h | Backup econômico |
| | MiniMax M2.1 | $0.2/1M | Rotativo 5h | Opção mais barata |
| | Kimi K2 | $9/mês fixo | 10M tokens/mês | Custo previsível |
| **🆓 GRATUITO** | iFlow | $0 | Ilimitado | 8 modelos gratuitos |
| | Qwen | $0 | Ilimitado | 3 modelos gratuitos |
| | Kiro | $0 | Ilimitado | Claude gratuito |
| Tier | Provedor | Custo | Reset de Cota | Melhor Para |
| ----------------- | ----------------- | ---------------------------- | ----------------- | ------------------------------ |
| **💳 ASSINATURA** | Claude Code (Pro) | $20/mês | 5h + semanal | Já é assinante |
| | Codex (Plus/Pro) | $20-200/mês | 5h + semanal | Usuários OpenAI |
| | Gemini CLI | **GRATUITO** | 180K/mês + 1K/dia | Todos! |
| | GitHub Copilot | $10-19/mês | Mensal | Usuários GitHub |
| **🔑 API KEY** | NVIDIA NIM | **GRATUITO** (1000 créditos) | Único | Testes gratuitos |
| | DeepSeek | Por uso | Nenhum | Melhor preço/qualidade |
| | Groq | Tier gratuito + pago | Limitado | Inferência ultra-rápida |
| | xAI (Grok) | Por uso | Nenhum | Modelos Grok |
| | Mistral | Tier gratuito + pago | Limitado | IA Europeia |
| | OpenRouter | Por uso | Nenhum | 100+ modelos |
| **💰 BARATO** | GLM-4.7 | $0.6/1M | Diário 10h | Backup econômico |
| | MiniMax M2.1 | $0.2/1M | Rotativo 5h | Opção mais barata |
| | Kimi K2 | $9/mês fixo | 10M tokens/mês | Custo previsível |
| **🆓 GRATUITO** | iFlow | $0 | Ilimitado | 8 modelos gratuitos |
| | Qwen | $0 | Ilimitado | 3 modelos gratuitos |
| | Kiro | $0 | Ilimitado | Claude gratuito |
| | LongCat 🆕 | **$0** (50M tok/dia 🔥) | 1 req/s | Maior cota grátis do mundo |
| | Pollinations 🆕 | **$0** (sem chave API) | 1 req/15s | GPT-5, Claude, DeepSeek, Llama |
| | Cloudflare AI 🆕 | **$0** (10K Neurons/dia) | ~150 resp/dia | 50+ modelos, edge global |
| | Scaleway AI 🆕 | **$0** (1M tokens total) | Limitado por taxa | EU/GDPR, Qwen3 235B, Llama 70B |
**💡 Dica Pro:** Comece com Gemini CLI (180K grátis/mês) + iFlow (ilimitado grátis) = $0 de custo!
@@ -879,16 +883,16 @@ Por que isso é relevante:
### 🎵 APIs Multi-Modal
| Funcionalidade | O que Faz |
| --------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 🖼️ **Geração de Imagem** | `/v1/images/generations` — 10 provedores, 20+ modelos (cloud + local) |
| 📐 **Embeddings** | `/v1/embeddings` — 6 provedores, 9+ modelos |
| 🎤 **Transcrição de Áudio** | `/v1/audio/transcriptions`Whisper + Nvidia NIM, HuggingFace, Qwen3 |
| 🔊 **Texto para Fala** | `/v1/audio/speech`ElevenLabs, Nvidia NIM, HuggingFace, Coqui, Tortoise, Qwen3, Inworld, Cartesia, PlayHT |
| 🎬 **Geração de Vídeo** | `/v1/videos/generations` — ComfyUI (AnimateDiff, SVD), SD WebUI |
| 🎵 **Geração de Música** | `/v1/music/generations` — ComfyUI (Stable Audio Open, MusicGen) |
| 🛡️ **Moderações** | `/v1/moderations` — Verificações de segurança |
| 🔀 **Reranking** | `/v1/rerank` — Reranking de relevância de documentos |
| Funcionalidade | O que Faz |
| --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Geração de Imagem** | `/v1/images/generations` — 10 provedores, 20+ modelos (cloud + local) |
| 📐 **Embeddings** | `/v1/embeddings` — 6 provedores, 9+ modelos |
| 🎤 **Transcrição de Áudio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Texto para Fala** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🎬 **Geração de Vídeo** | `/v1/videos/generations` — ComfyUI (AnimateDiff, SVD), SD WebUI |
| 🎵 **Geração de Música** | `/v1/music/generations` — ComfyUI (Stable Audio Open, MusicGen) |
| 🛡️ **Moderações** | `/v1/moderations` — Verificações de segurança |
| 🔀 **Reranking** | `/v1/rerank` — Reranking de relevância de documentos |
### 🛡️ Resiliência e Segurança
@@ -1223,6 +1227,53 @@ Modelos:
kr/claude-haiku-4.5
```
### LongCat AI (GRATUITO 50M tokens/dia!) 🆕
1. Cadastre-se: [longcat.chat](https://longcat.chat) com e-mail ou telefone
2. Gere uma chave de API gratuita
3. Dashboard → Adicionar Provedor → LongCat
**Modelos:**
- `lc/LongCat-Flash-Lite`**50M tokens/dia** 💥 (maior cota gratuita do mundo!)
- `lc/LongCat-Flash-Chat` — 500K tokens/dia
- `lc/LongCat-Flash-Thinking` — 500K tokens/dia (raciocínio)
> 100% gratuito durante o beta público. Reset diário à meia-noite UTC.
### Pollinations AI (SEM CHAVE NECESSÁRIA!) 🆕
1. Adicione o provedor Pollinations no Dashboard
2. Deixe o campo de chave API vazio (ou coloque qualquer string)
3. Comece a usar imediatamente!
**Modelos via `pol/`:** `openai` (GPT-5), `claude`, `gemini`, `deepseek`, `llama` (Llama 4)
> Sem cadastro, sem chave, sem cartão de crédito. 1 req/15s ilimitado.
### Cloudflare Workers AI (GRATUITO 10K Neurons/dia!) 🆕
1. Cadastre-se: [dash.cloudflare.com](https://dash.cloudflare.com)
2. Gere um API Token em Profile → API Tokens
3. Copie seu Account ID (coluna direita do dashboard)
4. Dashboard → Adicionar Provedor → Cloudflare AI
- API Key: seu token
- Account ID: seu account ID
**Modelos via `cf/`:** `@cf/meta/llama-3.3-70b-instruct`, `@cf/google/gemma-3-12b-it`, 50+ mais
> 10K Neurons/dia ≈ 150 respostas de LLM ou 500s de transcrição Whisper gratuita!
### Scaleway AI (1M tokens gratuitos!) 🆕
1. Cadastre-se: [console.scaleway.com](https://console.scaleway.com)
2. Gere uma chave de API IAM
3. Dashboard → Adicionar Provedor → Scaleway
**Modelos via `scw/`:** `qwen3-235b-a22b-instruct-2507` (Qwen3 235B!), `llama-3.1-70b-instruct`
> 1M tokens gratuitos para novas contas. Dados processados na 🇫🇷 França (EU/GDPR).
</details>
<details>
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 APIs multimodais
| Recurso | O que faz |
| --------------------------------- | ----------------------------------------------------------- |
| 🖼️ **Geração de imagens** | `/v1/images/generations` — 4 provedores, mais de 9 modelos |
| 📐 **Incorporações** | `/v1/embeddings` — 6 provedores, mais de 9 modelos |
| 🎤 **Transcrição de áudio** | `/v1/audio/transcriptions`Compatível com sussurro |
| 🔊 **Conversão de texto em fala** | `/v1/audio/speech`Síntese de áudio multiprovedor |
| 🛡️ **Moderações** | `/v1/moderations` — Verificações de segurança de conteúdo |
| 🔀 **Reclassificação** | `/v1/rerank` — Reclassificação da relevância dos documentos |
| Recurso | O que faz |
| --------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Geração de imagens** | `/v1/images/generations` — 4 provedores, mais de 9 modelos |
| 📐 **Incorporações** | `/v1/embeddings` — 6 provedores, mais de 9 modelos |
| 🎤 **Transcrição de áudio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Conversão de texto em fala** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderações** | `/v1/moderations` — Verificações de segurança de conteúdo |
| 🔀 **Reclassificação** | `/v1/rerank` — Reclassificação da relevância dos documentos |
### 🛡️ Resiliência e segurança
+8 -8
View File
@@ -875,14 +875,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 API-uri multimodale
| Caracteristica | Ce face |
| ------------------------- | ---------------------------------------------------------- |
| 🖼️ **Generarea imaginii** | `/v1/images/generations` — 4 furnizori, peste 9 modele |
| 📐 **Inglobări** | `/v1/embeddings` — 6 furnizori, peste 9 modele |
| 🎤 **Transcriere audio** | `/v1/audio/transcriptions`Compatibil cu Whisper |
| 🔊 **Text-to-speech** | `/v1/audio/speech`Sinteză audio cu mai mulți furnizori |
| 🛡️ **Moderații** | `/v1/moderations` — Verificări de siguranță a conținutului |
| 🔀 **Reclasificare** | `/v1/rerank` — Reclasificarea relevanței documentului |
| Caracteristica | Ce face |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Generarea imaginii** | `/v1/images/generations` — 4 furnizori, peste 9 modele |
| 📐 **Inglobări** | `/v1/embeddings` — 6 furnizori, peste 9 modele |
| 🎤 **Transcriere audio** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Text-to-speech** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderații** | `/v1/moderations` — Verificări de siguranță a conținutului |
| 🔀 **Reclasificare** | `/v1/rerank` — Reclasificarea relevanței documentului |
### 🛡️ Reziliență și securitate
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Мультимодальные API
| Функция | Что делает |
| ---------------------------- | --------------------------------------------------- |
| 🖼️ **Генерация изображений** | `/v1/images/generations` — 4 провайдера, 9+ моделей |
| 📐 **Embeddings** | `/v1/embeddings` — 6 провайдеров, 9+ моделей |
| 🎤 **Транскрипция аудио** | `/v1/audio/transcriptions`Совместимо с Whisper |
| 🔊 **Текст в речь** | `/v1/audio/speech`Мульти-провайдерный синтез |
| 🛡️ **Модерация** | `/v1/moderations` — Проверки безопасности контента |
| 🔀 **Reranking** | `/v1/rerank` — Переранжирование релевантности |
| Функция | Что делает |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Генерация изображений** | `/v1/images/generations` — 4 провайдера, 9+ моделей |
| 📐 **Embeddings** | `/v1/embeddings` — 6 провайдеров, 9+ моделей |
| 🎤 **Транскрипция аудио** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Текст в речь** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Модерация** | `/v1/moderations` — Проверки безопасности контента |
| 🔀 **Reranking** | `/v1/rerank` — Переранжирование релевантности |
### 🛡️ Устойчивость и безопасность
+8 -8
View File
@@ -877,14 +877,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multimodálne API
| Funkcia | Čo to robí |
| --------------------------- | ---------------------------------------------------------------- |
| 🖼️ **Generovanie obrázkov** | `/v1/images/generations` — 4 poskytovatelia, 9+ modelov |
| 📐 **Vloženie** | `/v1/embeddings` — 6 poskytovateľov, 9+ modelov |
| 🎤 **Prepis zvuku** | `/v1/audio/transcriptions`Kompatibilné so šepotom |
| 🔊 **Prevod textu na reč** | `/v1/audio/speech`Zvuková syntéza od viacerých poskytovateľov |
| 🛡️ **Moderovania** | `/v1/moderations` — Kontroly bezpečnosti obsahu |
| 🔀 **Reranking** | `/v1/rerank` — Zmena poradia relevantnosti dokumentu |
| Funkcia | Čo to robí |
| --------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Generovanie obrázkov** | `/v1/images/generations` — 4 poskytovatelia, 9+ modelov |
| 📐 **Vloženie** | `/v1/embeddings` — 6 poskytovateľov, 9+ modelov |
| 🎤 **Prepis zvuku** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Prevod textu na reč** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderovania** | `/v1/moderations` — Kontroly bezpečnosti obsahu |
| 🔀 **Reranking** | `/v1/rerank` — Zmena poradia relevantnosti dokumentu |
### 🛡️ Odolnosť a bezpečnosť
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multimodala API:er
| Funktion | Vad det gör |
| ------------------------ | ------------------------------------------------------ |
| 🖼️ **Bildgenerering** | `/v1/images/generations` — 4 leverantörer, 9+ modeller |
| 📐 **Inbäddningar** | `/v1/embeddings` — 6 leverantörer, 9+ modeller |
| 🎤 **Ljudtranskription** | `/v1/audio/transcriptions`Whisper-kompatibel |
| 🔊 **Text-till-tal** | `/v1/audio/speech`Ljudsyntes med flera leverantörer |
| 🛡️ **Moderationer** | `/v1/moderations` — Innehållssäkerhetskontroller |
| 🔀 **Omrankning** | `/v1/rerank` — Omrankning av dokumentrelevans |
| Funktion | Vad det gör |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Bildgenerering** | `/v1/images/generations` — 4 leverantörer, 9+ modeller |
| 📐 **Inbäddningar** | `/v1/embeddings` — 6 leverantörer, 9+ modeller |
| 🎤 **Ljudtranskription** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Text-till-tal** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Moderationer** | `/v1/moderations` — Innehållssäkerhetskontroller |
| 🔀 **Omrankning** | `/v1/rerank` — Omrankning av dokumentrelevans |
### 🛡️ Motståndskraft och säkerhet
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Multi-Modal API
| คุณสมบัติ | มันทำอะไร |
| ----------------------- | ------------------------------------------------------------- |
| 🖼️ **การสร้างภาพ** | `/v1/images/generations` — ผู้ให้บริการ 4 ราย รุ่น 9+ |
| 📐 **การฝัง** | `/v1/embeddings` — ผู้ให้บริการ 6 ราย รุ่น 9+ |
| 🎶 **การถอดเสียง** | `/v1/audio/transcriptions` — รองรับการกระซิบ |
| 🔊 **ข้อความเป็นคำพูด** | `/v1/audio/speech`การสังเคราะห์เสียงจากผู้ให้บริการหลายราย |
| 🛡️ **การกลั่นกรอง** | `/v1/moderations` — การตรวจสอบความปลอดภัยของเนื้อหา |
| 🔀 **จัดอันดับ** | `/v1/rerank` — การจัดอันดับความเกี่ยวข้องของเอกสาร |
| คุณสมบัติ | มันทำอะไร |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **การสร้างภาพ** | `/v1/images/generations` — ผู้ให้บริการ 4 ราย รุ่น 9+ |
| 📐 **การฝัง** | `/v1/embeddings` — ผู้ให้บริการ 6 ราย รุ่น 9+ |
| 🎶 **การถอดเสียง** | `/v1/audio/transcriptions` — รองรับการกระซิบ |
| 🔊 **ข้อความเป็นคำพูด** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **การกลั่นกรอง** | `/v1/moderations` — การตรวจสอบความปลอดภัยของเนื้อหา |
| 🔀 **จัดอันดับ** | `/v1/rerank` — การจัดอันดับความเกี่ยวข้องของเอกสาร |
### 🛡️ ความยืดหยุ่นและความปลอดภัย
+8 -8
View File
@@ -878,14 +878,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 Мультимодальні API
| Особливість | Що він робить |
| ---------------------------------- | ----------------------------------------------------- |
| 🖼️ **Створення зображень** | `/v1/images/generations` — 4 провайдери, 9+ моделей |
| 📐 **Вбудовування** | `/v1/embeddings` — 6 провайдерів, 9+ моделей |
| 🎤 **Транскрипція аудіо** | `/v1/audio/transcriptions`сумісний із Whisper |
| 🔊 **Створення тексту в мовлення** | `/v1/audio/speech`Багатопровайдерний аудіосинтез |
| 🛡️ **Модерації** | `/v1/moderations` — Перевірка безпеки вмісту |
| 🔀 **Переранжування** | `/v1/rerank` — Переранжування релевантності документа |
| Особливість | Що він робить |
| ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Створення зображень** | `/v1/images/generations` — 4 провайдери, 9+ моделей |
| 📐 **Вбудовування** | `/v1/embeddings` — 6 провайдерів, 9+ моделей |
| 🎤 **Транскрипція аудіо** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Створення тексту в мовлення** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Модерації** | `/v1/moderations` — Перевірка безпеки вмісту |
| 🔀 **Переранжування** | `/v1/rerank` — Переранжування релевантності документа |
### 🛡️ Стійкість і безпека
+8 -8
View File
@@ -874,14 +874,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 API đa phương thức
| Tính năng | Nó làm gì |
| ------------------------------------- | ------------------------------------------------------------ |
| 🖼️ **Tạo hình ảnh** | `/v1/images/generations` — 4 nhà cung cấp, hơn 9 mô hình |
| 📐 **Nhúng** | `/v1/embeddings` — 6 nhà cung cấp, hơn 9 mô hình |
| 🎤 **Phiên âm âm thanh** | `/v1/audio/transcriptions`Tương thích với lời thì thầm |
| 🔊 **Chuyển văn bản thành giọng nói** | `/v1/audio/speech`Tổng hợp âm thanh từ nhiều nhà cung cấp |
| 🛡️ **Kiểm duyệt** | `/v1/moderations` — Kiểm tra an toàn nội dung |
| 🔀 **Sắp xếp lại** | `/v1/rerank` — Sắp xếp lại mức độ liên quan của tài liệu |
| Tính năng | Nó làm gì |
| ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **Tạo hình ảnh** | `/v1/images/generations` — 4 nhà cung cấp, hơn 9 mô hình |
| 📐 **Nhúng** | `/v1/embeddings` — 6 nhà cung cấp, hơn 9 mô hình |
| 🎤 **Phiên âm âm thanh** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **Chuyển văn bản thành giọng nói** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **Kiểm duyệt** | `/v1/moderations` — Kiểm tra an toàn nội dung |
| 🔀 **Sắp xếp lại** | `/v1/rerank` — Sắp xếp lại mức độ liên quan của tài liệu |
### 🛡️ Khả năng phục hồi và bảo mật
+8 -8
View File
@@ -873,14 +873,14 @@ npm run electron:build:linux # Linux (.AppImage)
### 🎵 多模态 API
| 功能 | 功能描述 |
| ----------------- | ---------------------------------------------- |
| 🖼️ **图像生成** | `/v1/images/generations` — 4 个提供商,9+ 模型 |
| 📐 **Embeddings** | `/v1/embeddings` — 6 个提供商,9+ 模型 |
| 🎤 **音频转录** | `/v1/audio/transcriptions`Whisper 兼容 |
| 🔊 **文字转语音** | `/v1/audio/speech`多提供商音频合成 |
| 🛡️ **内容审核** | `/v1/moderations` — 内容安全检查 |
| 🔀 **重排序** | `/v1/rerank` — 文档相关性重排序 |
| 功能 | 功能描述 |
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 🖼️ **图像生成** | `/v1/images/generations` — 4 个提供商,9+ 模型 |
| 📐 **Embeddings** | `/v1/embeddings` — 6 个提供商,9+ 模型 |
| 🎤 **音频转录** | `/v1/audio/transcriptions`7 providers (Deepgram Nova 3, AssemblyAI, Groq Whisper, HuggingFace, ElevenLabs, OpenAI, Azure), auto-language detection, MP4/MP3/WAV support |
| 🔊 **文字转语音** | `/v1/audio/speech`10 providers (ElevenLabs, OpenAI, Deepgram, Cartesia, PlayHT, HuggingFace, Nvidia NIM, Inworld, Coqui, Tortoise) |
| 🛡️ **内容审核** | `/v1/moderations` — 内容安全检查 |
| 🔀 **重排序** | `/v1/rerank` — 文档相关性重排序 |
### 🛡️ 弹性与安全
+1 -1
View File
@@ -1,7 +1,7 @@
openapi: 3.1.0
info:
title: OmniRoute API
version: 2.8.7
version: 3.0.0-rc.7
description: |
OmniRoute is a local-first AI API proxy router. It provides an OpenAI-compatible
endpoint that routes requests to multiple AI providers with load balancing,
+1
View File
@@ -17,6 +17,7 @@ export interface EmbeddingProvider {
}
export interface EmbeddingProviderNodeRow {
id?: string;
prefix: string;
name: string;
baseUrl: string;
+236
View File
@@ -47,6 +47,8 @@ export interface RegistryEntry {
executor: string;
baseUrl?: string;
baseUrls?: string[];
/** Override base URL used only for API key validation (e.g., opencode-go validates on zen/v1) */
testKeyBaseUrl?: string;
responsesBaseUrl?: string;
urlSuffix?: string;
urlBuilder?: (base: string, model: string, stream: boolean) => string;
@@ -495,6 +497,41 @@ export const REGISTRY: Record<string, RegistryEntry> = {
],
},
"opencode-go": {
id: "opencode-go",
alias: "opencode-go",
format: "openai",
executor: "opencode",
baseUrl: "https://opencode.ai/zen/go/v1",
// (#532) Key validation must hit the main zen endpoint (same key works for both tiers)
testKeyBaseUrl: "https://opencode.ai/zen/v1",
authType: "apikey",
authHeader: "Authorization",
authPrefix: "Bearer",
models: [
{ id: "glm-5", name: "GLM-5" },
{ id: "kimi-k2.5", name: "Kimi K2.5" },
{ id: "minimax-m2.7", name: "MiniMax M2.7", targetFormat: "claude" },
{ id: "minimax-m2.5", name: "MiniMax M2.5", targetFormat: "claude" },
],
},
"opencode-zen": {
id: "opencode-zen",
alias: "opencode-zen",
format: "openai",
executor: "opencode",
baseUrl: "https://opencode.ai/zen/v1",
authType: "apikey",
authHeader: "Authorization",
authPrefix: "Bearer",
models: [
{ id: "minimax-m2.5-free", name: "MiniMax M2.5 Free" },
{ id: "big-pickle", name: "Big Pickle" },
{ id: "gpt-5-nano", name: "GPT 5 Nano" },
],
},
openrouter: {
id: "openrouter",
alias: "openrouter",
@@ -883,6 +920,12 @@ export const REGISTRY: Record<string, RegistryEntry> = {
authType: "apikey",
authHeader: "bearer",
models: [
{ id: "meta-llama/Llama-3.3-70B-Instruct-Turbo-Free", name: "Llama 3.3 70B Turbo (🆓 Free)" },
{ id: "meta-llama/Llama-Vision-Free", name: "Llama Vision (🆓 Free)" },
{
id: "deepseek-ai/DeepSeek-R1-Distill-Llama-70B-Free",
name: "DeepSeek R1 Distill 70B (🆓 Free)",
},
{ id: "meta-llama/Llama-3.3-70B-Instruct-Turbo", name: "Llama 3.3 70B Turbo" },
{ id: "deepseek-ai/DeepSeek-R1", name: "DeepSeek R1" },
{ id: "Qwen/Qwen3-235B-A22B", name: "Qwen3 235B" },
@@ -1125,6 +1168,199 @@ export const REGISTRY: Record<string, RegistryEntry> = {
{ id: "claude-sonnet-4-5@20251101", name: "Claude Sonnet 4.5 (Vertex)" },
],
},
alibaba: {
id: "alibaba",
alias: "ali",
format: "openai",
executor: "default",
// DashScope international OpenAI-compatible endpoint.
// China users should set providerSpecificData.baseUrl to:
// https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions
baseUrl: "https://dashscope-intl.aliyuncs.com/compatible-mode/v1/chat/completions",
modelsUrl: "https://dashscope-intl.aliyuncs.com/compatible-mode/v1/models",
authType: "apikey",
authHeader: "bearer",
models: [
{ id: "qwen-max", name: "Qwen Max" },
{ id: "qwen-max-2025-01-25", name: "Qwen Max (2025-01-25)" },
{ id: "qwen-plus", name: "Qwen Plus" },
{ id: "qwen-plus-2025-07-14", name: "Qwen Plus (2025-07-14)" },
{ id: "qwen-turbo", name: "Qwen Turbo" },
{ id: "qwen-turbo-2025-11-01", name: "Qwen Turbo (2025-11-01)" },
{ id: "qwen3-coder-plus", name: "Qwen3 Coder Plus" },
{ id: "qwen3-coder-flash", name: "Qwen3 Coder Flash" },
{ id: "qwq-plus", name: "QwQ Plus (Reasoning)" },
{ id: "qwq-32b", name: "QwQ 32B" },
{ id: "qwen3-32b", name: "Qwen3 32B" },
{ id: "qwen3-235b-a22b", name: "Qwen3 235B A22B" },
],
passthroughModels: true,
},
// ── New Free Providers (2026) ─────────────────────────────────────────────
longcat: {
id: "longcat",
alias: "lc",
format: "openai",
executor: "default",
// (#536) Correct OpenAI-compatible base URL — was longcat.chat/api/v1/chat/completions
// which is the chat endpoint directly, not the base. Key validation and routing must
// use https://api.longcat.chat/openai which resolves /v1/models and /v1/chat/completions
baseUrl: "https://api.longcat.chat/openai",
authType: "apikey",
authHeader: "Authorization",
authPrefix: "Bearer",
// Free tier: 50M tokens/day (Flash-Lite) + 500K/day (Chat/Thinking) — 100% free while public beta
models: [
{ id: "LongCat-Flash-Lite", name: "LongCat Flash-Lite (50M tok/day 🆓)" },
{ id: "LongCat-Flash-Chat", name: "LongCat Flash-Chat (500K tok/day 🆓)" },
{ id: "LongCat-Flash-Thinking", name: "LongCat Flash-Thinking (500K tok/day 🆓)" },
{ id: "LongCat-Flash-Thinking-2601", name: "LongCat Flash-Thinking-2601 (🆓)" },
{ id: "LongCat-Flash-Omni-2603", name: "LongCat Flash-Omni-2603 (🆓)" },
],
},
pollinations: {
id: "pollinations",
alias: "pol",
format: "openai",
executor: "pollinations",
// No API key required for basic use. Proxy to GPT-5, Claude, Gemini, DeepSeek, Llama 4.
baseUrl: "https://text.pollinations.ai/openai/chat/completions",
authType: "apikey", // Optional — works without one too
authHeader: "bearer",
models: [
{ id: "openai", name: "GPT-5 via Pollinations (🆓)" },
{ id: "claude", name: "Claude via Pollinations (🆓)" },
{ id: "gemini", name: "Gemini via Pollinations (🆓)" },
{ id: "deepseek", name: "DeepSeek V3 via Pollinations (🆓)" },
{ id: "llama", name: "Llama 4 via Pollinations (🆓)" },
{ id: "mistral", name: "Mistral via Pollinations (🆓)" },
],
},
puter: {
id: "puter",
alias: "pu",
format: "openai",
executor: "puter",
// OpenAI-compatible gateway with 500+ models (GPT, Claude, Gemini, Grok, DeepSeek, Qwen…)
// Auth: Bearer <puter_auth_token> from puter.com/dashboard → Copy Auth Token
// Model IDs use provider/model-name format for non-OpenAI models.
// Only chat completions (incl. streaming) are available via REST.
// Image gen, TTS, STT, video are puter.js SDK-only (browser).
baseUrl: "https://api.puter.com/puterai/openai/v1/chat/completions",
authType: "apikey",
authHeader: "bearer",
models: [
// OpenAI — use bare IDs
{ id: "gpt-4o-mini", name: "GPT-4o Mini (🆓 Puter)" },
{ id: "gpt-4o", name: "GPT-4o (Puter)" },
{ id: "gpt-4.1", name: "GPT-4.1 (Puter)" },
{ id: "gpt-4.1-mini", name: "GPT-4.1 Mini (Puter)" },
{ id: "gpt-5-nano", name: "GPT-5 Nano (Puter)" },
{ id: "gpt-5-mini", name: "GPT-5 Mini (Puter)" },
{ id: "gpt-5", name: "GPT-5 (Puter)" },
{ id: "o3-mini", name: "OpenAI o3-mini (Puter)" },
{ id: "o3", name: "OpenAI o3 (Puter)" },
{ id: "o4-mini", name: "OpenAI o4-mini (Puter)" },
// Anthropic Claude — use bare IDs (confirmed working)
{ id: "claude-haiku-4-5", name: "Claude Haiku 4.5 (Puter)" },
{ id: "claude-sonnet-4-5", name: "Claude Sonnet 4.5 (Puter)" },
{ id: "claude-opus-4-5", name: "Claude Opus 4.5 (Puter)" },
{ id: "claude-sonnet-4", name: "Claude Sonnet 4 (Puter)" },
{ id: "claude-opus-4", name: "Claude Opus 4 (Puter)" },
// Google Gemini — use google/ prefix (confirmed working)
{ id: "google/gemini-2.0-flash", name: "Gemini 2.0 Flash (Puter)" },
{ id: "google/gemini-2.5-flash", name: "Gemini 2.5 Flash (Puter)" },
{ id: "google/gemini-2.5-pro", name: "Gemini 2.5 Pro (Puter)" },
{ id: "google/gemini-3-flash", name: "Gemini 3 Flash (Puter)" },
{ id: "google/gemini-3-pro", name: "Gemini 3 Pro (Puter)" },
// DeepSeek — use deepseek/ prefix (confirmed working)
{ id: "deepseek/deepseek-chat", name: "DeepSeek Chat (Puter)" },
{ id: "deepseek/deepseek-r1", name: "DeepSeek R1 (Puter)" },
{ id: "deepseek/deepseek-v3.2", name: "DeepSeek V3.2 (Puter)" },
// xAI Grok — use x-ai/ prefix
{ id: "x-ai/grok-3", name: "Grok 3 (Puter)" },
{ id: "x-ai/grok-3-mini", name: "Grok 3 Mini (Puter)" },
{ id: "x-ai/grok-4", name: "Grok 4 (Puter)" },
{ id: "x-ai/grok-4-fast", name: "Grok 4 Fast (Puter)" },
// Meta Llama — bare IDs (confirmed ✅)
{ id: "llama-4-scout", name: "Llama 4 Scout (Puter)" },
{ id: "llama-4-maverick", name: "Llama 4 Maverick (Puter)" },
{ id: "llama-3.3-70b-instruct", name: "Llama 3.3 70B (Puter)" },
// Mistral — bare IDs (confirmed ✅)
{ id: "mistral-small-latest", name: "Mistral Small (Puter)" },
{ id: "mistral-medium-latest", name: "Mistral Medium (Puter)" },
{ id: "open-mistral-nemo", name: "Mistral Nemo (Puter)" },
// Qwen — use qwen/ prefix (confirmed ✅)
{ id: "qwen/qwen3-235b-a22b", name: "Qwen3 235B (Puter)" },
{ id: "qwen/qwen3-32b", name: "Qwen3 32B (Puter)" },
{ id: "qwen/qwen3-coder", name: "Qwen3 Coder 480B (Puter)" },
],
passthroughModels: true, // 500+ models available — users can type any Puter model ID
},
"cloudflare-ai": {
id: "cloudflare-ai",
alias: "cf",
format: "openai",
executor: "cloudflare-ai",
// URL is dynamic: uses accountId from credentials. The executor builds it.
baseUrl: "https://api.cloudflare.com/client/v4/accounts",
authType: "apikey",
authHeader: "bearer",
// 10K Neurons/day free: ~150 LLM responses or 500s Whisper audio — global edge
models: [
{ id: "@cf/meta/llama-3.3-70b-instruct", name: "Llama 3.3 70B (🆓 ~150 resp/day)" },
{ id: "@cf/meta/llama-3.1-8b-instruct", name: "Llama 3.1 8B (🆓)" },
{ id: "@cf/google/gemma-3-12b-it", name: "Gemma 3 12B (🆓)" },
{ id: "@cf/mistral/mistral-7b-instruct-v0.2-lora", name: "Mistral 7B (🆓)" },
{ id: "@cf/qwen/qwen2.5-coder-15b-instruct", name: "Qwen 2.5 Coder 15B (🆓)" },
{ id: "@cf/deepseek-ai/deepseek-r1-distill-qwen-32b", name: "DeepSeek R1 Distill 32B (🆓)" },
],
},
scaleway: {
id: "scaleway",
alias: "scw",
format: "openai",
executor: "default",
baseUrl: "https://api.scaleway.ai/v1/chat/completions",
authType: "apikey",
authHeader: "bearer",
// 1M tokens free for new accounts — EU/GDPR (Paris), no credit card needed under limit
models: [
{ id: "qwen3-235b-a22b-instruct-2507", name: "Qwen3 235B A22B (1M free tok 🆓)" },
{ id: "llama-3.1-70b-instruct", name: "Llama 3.1 70B (🆓 EU)" },
{ id: "llama-3.1-8b-instruct", name: "Llama 3.1 8B (🆓 EU)" },
{ id: "mistral-small-3.2-24b-instruct-2506", name: "Mistral Small 3.2 (🆓 EU)" },
{ id: "deepseek-v3-0324", name: "DeepSeek V3 (🆓 EU)" },
{ id: "gpt-oss-120b", name: "GPT-OSS 120B (🆓 EU)" },
],
},
aimlapi: {
id: "aimlapi",
alias: "aiml",
format: "openai",
executor: "default",
baseUrl: "https://api.aimlapi.com/v1/chat/completions",
authType: "apikey",
authHeader: "bearer",
// $0.025/day free credits — 200+ models via single aggregator endpoint
models: [
{ id: "gpt-4o", name: "GPT-4o (via AI/ML API)" },
{ id: "claude-3-5-sonnet-20241022", name: "Claude 3.5 Sonnet (via AI/ML API)" },
{ id: "gemini-1.5-pro", name: "Gemini 1.5 Pro (via AI/ML API)" },
{ id: "meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo", name: "Llama 3.1 70B (via AI/ML API)" },
{ id: "deepseek-chat", name: "DeepSeek Chat (via AI/ML API)" },
{ id: "mistral-large-latest", name: "Mistral Large (via AI/ML API)" },
],
passthroughModels: true,
},
};
// ── Generator Functions ───────────────────────────────────────────────────
+20 -4
View File
@@ -44,12 +44,28 @@ export class AntigravityExecutor extends BaseExecutor {
// stale/wrong client-side values causing 404/403 from Cloud Code endpoints.
// Opt-in escape hatch: set OMNIROUTE_ALLOW_BODY_PROJECT_OVERRIDE=1.
const projectId =
allowBodyProjectOverride && bodyProjectId ? bodyProjectId : credentialsProjectId || bodyProjectId;
allowBodyProjectOverride && bodyProjectId
? bodyProjectId
: credentialsProjectId || bodyProjectId;
if (!projectId) {
throw new Error(
"Missing Google projectId for Antigravity account. Please reconnect OAuth so OmniRoute can fetch your real Cloud Code project (loadCodeAssist)."
);
// (#489) Return a structured error instead of throwing — gives the client a clear signal
// to show a "Reconnect OAuth" prompt rather than an opaque "Internal Server Error".
const errorMsg =
"Missing Google projectId for Antigravity account. Please reconnect OAuth in Providers → Antigravity so OmniRoute can fetch your Cloud Code project.";
const errorBody = {
error: {
message: errorMsg,
type: "oauth_missing_project_id",
code: "missing_project_id",
},
};
const resp = new Response(JSON.stringify(errorBody), {
status: 422,
headers: { "Content-Type": "application/json" },
});
// Returning a Response object signals the executor to stop and forward it
return resp as unknown as never;
}
// Fix contents for Claude models via Antigravity
+17 -1
View File
@@ -2,6 +2,20 @@ import { HTTP_STATUS, FETCH_TIMEOUT_MS } from "../config/constants.ts";
import { applyFingerprint, isCliCompatEnabled } from "../config/cliFingerprints.ts";
import { getRotatingApiKey } from "../services/apiKeyRotator.ts";
/**
* Sanitizes a custom API path to prevent path traversal attacks.
* Valid paths must start with '/', contain no '..' segments,
* no null bytes, and be reasonable in length.
*/
function sanitizePath(path: string): boolean {
if (typeof path !== "string") return false;
if (!path.startsWith("/")) return false;
if (path.includes("\0")) return false; // null byte
if (path.includes("..")) return false; // path traversal
if (path.length > 512) return false; // sanity limit
return true;
}
type JsonRecord = Record<string, unknown>;
export type ProviderConfig = {
@@ -103,7 +117,9 @@ export class BaseExecutor {
const psd = credentials?.providerSpecificData;
const baseUrl = typeof psd?.baseUrl === "string" ? psd.baseUrl : "https://api.openai.com/v1";
const normalized = baseUrl.replace(/\/$/, "");
const customPath = typeof psd?.chatPath === "string" && psd.chatPath ? psd.chatPath : null;
// Sanitize custom path: must start with '/', no path traversal, no null bytes
const rawPath = typeof psd?.chatPath === "string" && psd.chatPath ? psd.chatPath : null;
const customPath = rawPath && sanitizePath(rawPath) ? rawPath : null;
if (customPath) return `${normalized}${customPath}`;
const path = this.provider.includes("responses") ? "/responses" : "/chat/completions";
return `${normalized}${path}`;
+59
View File
@@ -0,0 +1,59 @@
import { BaseExecutor } from "./base.ts";
import { PROVIDERS } from "../config/constants.ts";
/**
* CloudflareAIExecutor handles dynamic URL construction with accountId.
* Cloudflare Workers AI uses the authenticated user's account ID in the URL.
*
* URL pattern: https://api.cloudflare.com/client/v4/accounts/{accountId}/ai/v1/chat/completions
* Auth: Bearer <API Token>
* Docs: https://developers.cloudflare.com/workers-ai/
*
* Free tier: 10,000 Neurons/day = ~150 LLM responses or 500s Whisper audio
* API Token: dash.cloudflare.com/profile/api-tokens
* Account ID: right sidebar of dash.cloudflare.com
*/
export class CloudflareAIExecutor extends BaseExecutor {
constructor() {
super("cloudflare-ai", PROVIDERS["cloudflare-ai"] || { format: "openai" });
}
buildUrl(_model: string, _stream: boolean, _urlIndex = 0, credentials: any = null): string {
// Account ID can be stored in providerSpecificData or at top level credentials
const accountId =
credentials?.providerSpecificData?.accountId ||
credentials?.accountId ||
process.env.CLOUDFLARE_ACCOUNT_ID;
if (!accountId) {
throw new Error(
"Cloudflare Workers AI requires an Account ID. " +
"Add it in provider settings under 'Account ID'. " +
"Find it at: https://dash.cloudflare.com (right sidebar)."
);
}
return `https://api.cloudflare.com/client/v4/accounts/${accountId}/ai/v1/chat/completions`;
}
buildHeaders(credentials: any, stream = true): Record<string, string> {
const headers: Record<string, string> = {
"Content-Type": "application/json",
Authorization: `Bearer ${credentials.apiKey || credentials.accessToken}`,
};
if (stream) {
headers["Accept"] = "text/event-stream";
}
return headers;
}
transformRequest(_model: string, body: any, _stream: boolean, _credentials: any): any {
// Cloudflare uses full model paths like @cf/meta/llama-3.3-70b-instruct
// No transformation needed — user sends the full Cloudflare model path.
return body;
}
}
export default CloudflareAIExecutor;
+106
View File
@@ -3,6 +3,112 @@ import { CODEX_DEFAULT_INSTRUCTIONS } from "../config/codexInstructions.ts";
import { PROVIDERS } from "../config/constants.ts";
import { refreshCodexToken } from "../services/tokenRefresh.ts";
// ─── T09: Codex vs Spark Scope-Aware Rate Limiting ────────────────────────
// Codex has two independent quota pools: "codex" (standard) and "spark" (premium).
// Exhausting one should NOT block requests to the other.
// Ref: sub2api PR #1129 (feat(openai): split codex spark rate limiting from codex)
/**
* Maps model name substrings to their rate-limit scope.
* Checked in order first match wins.
*/
const CODEX_SCOPE_PATTERNS: Array<{ pattern: string; scope: "codex" | "spark" }> = [
{ pattern: "codex-spark", scope: "spark" },
{ pattern: "spark", scope: "spark" },
{ pattern: "codex", scope: "codex" },
{ pattern: "gpt-5", scope: "codex" }, // gpt-5.2-codex, gpt-5.3-codex, etc.
];
/**
* T09: Determine the rate-limit scope for a Codex model.
* Use this key as the suffix for per-scope rate limit state:
* `${accountId}:${getModelScope(model)}`
*
* @param model - The Codex model ID (e.g. "gpt-5.3-codex", "codex-spark-mini")
* @returns "codex" | "spark"
*/
export function getCodexModelScope(model: string): "codex" | "spark" {
const lower = model.toLowerCase();
for (const { pattern, scope } of CODEX_SCOPE_PATTERNS) {
if (lower.includes(pattern)) return scope;
}
return "codex"; // default scope
}
/**
* T09: Get the scope-keyed rate limit identifier for an account+model combination.
* Use this as the key for rateLimitState maps to ensure scope isolation.
*/
export function getCodexRateLimitKey(accountId: string, model: string): string {
return `${accountId}:${getCodexModelScope(model)}`;
}
/**
* T03: Parsed quota snapshot from Codex response headers.
* Codex includes per-account usage windows that allow precise reset scheduling.
* Ref: sub2api PR #357 (feat(oauth): persist usage snapshots and window cooldown)
*/
export interface CodexQuotaSnapshot {
usage5h: number; // tokens used in 5h window
limit5h: number; // token limit for 5h window
resetAt5h: string | null; // ISO timestamp when 5h window resets
usage7d: number; // tokens used in 7d window
limit7d: number; // token limit for 7d window
resetAt7d: string | null; // ISO timestamp when 7d window resets
}
/**
* T03: Parse Codex-specific quota headers from a provider response.
* Returns null if none of the relevant headers are present.
*
* Extracts:
* x-codex-5h-usage / x-codex-5h-limit / x-codex-5h-reset-at
* x-codex-7d-usage / x-codex-7d-limit / x-codex-7d-reset-at
*/
export function parseCodexQuotaHeaders(headers: Headers): CodexQuotaSnapshot | null {
const usage5h = headers.get("x-codex-5h-usage");
const limit5h = headers.get("x-codex-5h-limit");
const resetAt5h = headers.get("x-codex-5h-reset-at");
const usage7d = headers.get("x-codex-7d-usage");
const limit7d = headers.get("x-codex-7d-limit");
const resetAt7d = headers.get("x-codex-7d-reset-at");
// Return null if none of the quota headers are present (not a quota-aware response)
if (!usage5h && !limit5h && !resetAt5h && !usage7d && !limit7d && !resetAt7d) {
return null;
}
return {
usage5h: usage5h ? parseFloat(usage5h) : 0,
limit5h: limit5h ? parseFloat(limit5h) : Infinity,
resetAt5h: resetAt5h ?? null,
usage7d: usage7d ? parseFloat(usage7d) : 0,
limit7d: limit7d ? parseFloat(limit7d) : Infinity,
resetAt7d: resetAt7d ?? null,
};
}
/**
* T03: Get the soonest quota reset time from a CodexQuotaSnapshot.
* 7d window takes priority (wider window, harder limit) but we use whichever
* is further in the future to avoid releasing the block too early.
*
* @returns Unix timestamp (ms) of the soonest effective reset, or null
*/
export function getCodexResetTime(quota: CodexQuotaSnapshot): number | null {
const times: number[] = [];
if (quota.resetAt7d) {
const t = new Date(quota.resetAt7d).getTime();
if (!isNaN(t) && t > Date.now()) times.push(t);
}
if (quota.resetAt5h) {
const t = new Date(quota.resetAt5h).getTime();
if (!isNaN(t) && t > Date.now()) times.push(t);
}
if (times.length === 0) return null;
return Math.max(...times); // Use furthest-out reset to avoid premature unblock
}
// Ordered list of effort levels from lowest to highest
const EFFORT_ORDER = ["none", "low", "medium", "high", "xhigh"] as const;
type EffortLevel = (typeof EFFORT_ORDER)[number];
+6 -10
View File
@@ -80,18 +80,14 @@ export class DefaultExecutor extends BaseExecutor {
}
/**
* For compatible providers, ensure the model name sent upstream
* is the clean model name without internal routing prefixes.
* e.g. "openapi-chat-anti/claude-opus-4-6-thinking" "claude-opus-4-6-thinking"
* For compatible providers, the model name is already clean by the time
* it reaches the executor (chatCore sets body.model = modelInfo.model,
* which is the parsed model ID without internal routing prefixes).
*
* Models may legitimately contain "/" as part of their ID (e.g. "zai-org/GLM-5-FP8",
* "org/model-name") we must NOT strip path segments. (Fix #493)
*/
transformRequest(model, body, stream, credentials) {
if (
this.provider?.startsWith?.("openai-compatible-") ||
this.provider?.startsWith?.("anthropic-compatible-")
) {
const cleanModel = model.includes("/") ? model.split("/").slice(1).join("/") : model;
return { ...body, model: cleanModel };
}
return body;
}
+16
View File
@@ -6,6 +6,10 @@ import { KiroExecutor } from "./kiro.ts";
import { CodexExecutor } from "./codex.ts";
import { CursorExecutor } from "./cursor.ts";
import { DefaultExecutor } from "./default.ts";
import { PollinationsExecutor } from "./pollinations.ts";
import { CloudflareAIExecutor } from "./cloudflare-ai.ts";
import { OpencodeExecutor } from "./opencode.ts";
import { PuterExecutor } from "./puter.ts";
const executors = {
antigravity: new AntigravityExecutor(),
@@ -16,6 +20,14 @@ const executors = {
codex: new CodexExecutor(),
cursor: new CursorExecutor(),
cu: new CursorExecutor(), // Alias for cursor
pollinations: new PollinationsExecutor(),
pol: new PollinationsExecutor(), // Alias
"cloudflare-ai": new CloudflareAIExecutor(),
cf: new CloudflareAIExecutor(), // Alias
"opencode-zen": new OpencodeExecutor("opencode-zen"),
"opencode-go": new OpencodeExecutor("opencode-go"),
puter: new PuterExecutor(),
pu: new PuterExecutor(), // Alias
};
const defaultCache = new Map();
@@ -39,3 +51,7 @@ export { KiroExecutor } from "./kiro.ts";
export { CodexExecutor } from "./codex.ts";
export { CursorExecutor } from "./cursor.ts";
export { DefaultExecutor } from "./default.ts";
export { PollinationsExecutor } from "./pollinations.ts";
export { CloudflareAIExecutor } from "./cloudflare-ai.ts";
export { OpencodeExecutor } from "./opencode.ts";
export { PuterExecutor } from "./puter.ts";
+61
View File
@@ -0,0 +1,61 @@
import { BaseExecutor, type ExecuteInput, type ProviderCredentials } from "./base.ts";
import { PROVIDERS } from "../config/constants.ts";
import { getModelTargetFormat } from "../config/providerModels.ts";
export class OpencodeExecutor extends BaseExecutor {
_requestFormat: string | null = null;
constructor(provider: string) {
super(provider, PROVIDERS[provider] || PROVIDERS.openai);
}
async execute(input: ExecuteInput) {
this._requestFormat = getModelTargetFormat(this.provider, input.model) || "openai";
try {
return await super.execute(input);
} finally {
this._requestFormat = null;
}
}
buildUrl(
model: string,
stream: boolean,
urlIndex = 0,
credentials: ProviderCredentials | null = null
) {
void urlIndex;
void credentials;
const base = this.config.baseUrl;
switch (this._requestFormat) {
case "claude":
return `${base}/messages`;
case "openai-responses":
return `${base}/responses`;
case "gemini":
return `${base}/models/${model}:${stream ? "streamGenerateContent?alt=sse" : "generateContent"}`;
default:
return `${base}/chat/completions`;
}
}
buildHeaders(credentials: ProviderCredentials | null, stream = true) {
const headers: Record<string, string> = { "Content-Type": "application/json" };
const key = credentials?.apiKey || credentials?.accessToken;
if (key) {
headers["Authorization"] = `Bearer ${key}`;
}
if (this._requestFormat === "claude") {
headers["anthropic-version"] = "2023-06-01";
}
if (stream) {
headers["Accept"] = "text/event-stream";
}
return headers;
}
}
+46
View File
@@ -0,0 +1,46 @@
import { BaseExecutor } from "./base.ts";
import { PROVIDERS } from "../config/constants.ts";
/**
* PollinationsExecutor handles optional API key auth.
* Pollinations AI works WITHOUT any API key for basic use (1 req/15s).
* If an API key is provided, higher rate limits apply.
*
* Endpoint: https://text.pollinations.ai/openai/chat/completions
* Docs: https://pollinations.ai/docs
*/
export class PollinationsExecutor extends BaseExecutor {
constructor() {
super("pollinations", PROVIDERS["pollinations"] || { format: "openai" });
}
buildUrl(_model: string, _stream: boolean, _urlIndex = 0, _credentials = null): string {
return "https://text.pollinations.ai/openai/chat/completions";
}
buildHeaders(credentials: any, stream = true): Record<string, string> {
const headers: Record<string, string> = {
"Content-Type": "application/json",
};
// API key is OPTIONAL — skip Authorization header if no key provided
const key = credentials?.apiKey || credentials?.accessToken;
if (key) {
headers["Authorization"] = `Bearer ${key}`;
}
if (stream) {
headers["Accept"] = "text/event-stream";
}
return headers;
}
transformRequest(model: string, body: any, _stream: boolean, _credentials: any): any {
// Pollinations uses model names directly like "openai", "claude", "deepseek", etc.
// No transformation needed — the model name is already the Pollinations alias.
return body;
}
}
export default PollinationsExecutor;
+59
View File
@@ -0,0 +1,59 @@
import { BaseExecutor } from "./base.ts";
import { PROVIDERS } from "../config/constants.ts";
/**
* PuterExecutor OpenAI-compatible proxy for Puter AI.
*
* Puter exposes 500+ models (GPT, Claude, Gemini, Grok, DeepSeek, Qwen, Mistral...)
* through a single OpenAI-compatible REST endpoint.
*
* Endpoint: https://api.puter.com/puterai/openai/v1/chat/completions
* Auth: Bearer <puter_auth_token> (from puter.com/dashboard Copy Auth Token)
* Docs: https://docs.puter.com/AI/
*
* Model ID examples:
* OpenAI: "gpt-4o-mini", "gpt-4o", "gpt-4.1"
* Claude: "claude-sonnet-4-5", "claude-opus-4", "claude-haiku-4-5"
* Gemini: "google/gemini-2.0-flash", "google/gemini-2.5-pro"
* DeepSeek: "deepseek/deepseek-chat", "deepseek/deepseek-r1"
* Grok: "x-ai/grok-3", "x-ai/grok-4"
* Mistral: "mistralai/mistral-small-3.2"
* Meta: "meta-llama/llama-3.3-70b-instruct"
*
* Note: Image generation, TTS, STT, and video are puter.js SDK-only features.
* Only text chat completions (with streaming SSE) are available via REST.
*/
export class PuterExecutor extends BaseExecutor {
constructor() {
super("puter", PROVIDERS["puter"] || { format: "openai" });
}
buildUrl(_model: string, _stream: boolean, _urlIndex = 0, _credentials = null): string {
return "https://api.puter.com/puterai/openai/v1/chat/completions";
}
buildHeaders(credentials: any, stream = true): Record<string, string> {
const headers: Record<string, string> = {
"Content-Type": "application/json",
};
const key = credentials?.apiKey || credentials?.accessToken;
if (key) {
headers["Authorization"] = `Bearer ${key}`;
}
if (stream) {
headers["Accept"] = "text/event-stream";
}
return headers;
}
transformRequest(model: string, body: any, _stream: boolean, _credentials: any): any {
// Puter accepts model IDs directly from its catalog.
// No transformation required — model string is passed as-is.
return body;
}
}
export default PuterExecutor;
+8 -4
View File
@@ -28,13 +28,17 @@ function upstreamErrorResponse(res, errText) {
let errorMessage: string;
try {
const parsed = JSON.parse(errText);
errorMessage =
// Extract a human-readable message from various error response shapes.
// Guard against `parsed.error` being an object (e.g. ElevenLabs returns
// { error: { message: "...", status_code: 401 } } or { detail: { ... } })
const raw =
parsed?.err_msg ||
parsed?.error?.message ||
parsed?.error ||
(typeof parsed?.error === "string" ? parsed.error : null) ||
parsed?.message ||
parsed?.detail ||
errText;
(typeof parsed?.detail === "string" ? parsed.detail : parsed?.detail?.message) ||
null;
errorMessage = raw ? String(raw) : errText || `Upstream error (${res.status})`;
} catch {
errorMessage = errText || `Upstream error (${res.status})`;
}
+64 -8
View File
@@ -34,13 +34,15 @@ function upstreamErrorResponse(res, errText) {
let errorMessage: string;
try {
const parsed = JSON.parse(errText);
errorMessage =
// Guard against `parsed.error` or `parsed.detail` being objects
const raw =
parsed?.err_msg ||
parsed?.error?.message ||
parsed?.error ||
(typeof parsed?.error === "string" ? parsed.error : null) ||
parsed?.message ||
parsed?.detail ||
errText;
(typeof parsed?.detail === "string" ? parsed.detail : parsed?.detail?.message) ||
null;
errorMessage = raw ? String(raw) : errText || `Upstream error (${res.status})`;
} catch {
errorMessage = errText || `Upstream error (${res.status})`;
}
@@ -65,13 +67,67 @@ function getUploadedFileName(file: Blob & { name?: unknown }): string {
return typeof file.name === "string" && file.name.length > 0 ? file.name : "audio.wav";
}
/**
* Infer a suitable Content-Type for Deepgram from the browser-provided MIME
* type and the original filename. Deepgram accepts `audio/*` and many raw
* formats, but `video/*` causes it to silently fail with "no speech detected".
*
* Strategy:
* 1. If the browser says `audio/*`, keep it as-is.
* 2. If it's `video/*` (e.g. `.mp4`), remap to the audio equivalent so
* Deepgram extracts the audio track. `.mp4` `audio/mp4`, etc.
* 3. Fall back to `application/octet-stream` which tells Deepgram to
* auto-detect from the raw bytes (most reliable for unknown formats).
*/
function resolveAudioContentType(file: Blob & { name?: unknown }): string {
const browserType = (file.type || "").toLowerCase();
const fileName = typeof file.name === "string" ? file.name.toLowerCase() : "";
// 1) Browser already says it's audio — trust it
if (browserType.startsWith("audio/")) return browserType;
// 2) Derive from file extension (covers video/* and empty MIME)
const ext = fileName.includes(".") ? fileName.split(".").pop() : "";
const EXT_TO_MIME: Record<string, string> = {
mp3: "audio/mpeg",
mp4: "audio/mp4",
m4a: "audio/mp4",
wav: "audio/wav",
ogg: "audio/ogg",
flac: "audio/flac",
webm: "audio/webm",
aac: "audio/aac",
wma: "audio/x-ms-wma",
opus: "audio/opus",
};
if (ext && EXT_TO_MIME[ext]) return EXT_TO_MIME[ext];
// 3) Fallback — let Deepgram auto-detect from raw bytes
return "application/octet-stream";
}
/**
* Handle Deepgram transcription (raw binary audio, model via query param)
*/
async function handleDeepgramTranscription(providerConfig, file, modelId, token) {
async function handleDeepgramTranscription(
providerConfig,
file,
modelId,
token,
formData?: FormData
) {
const url = new URL(providerConfig.baseUrl);
url.searchParams.set("model", modelId);
url.searchParams.set("smart_format", "true");
url.searchParams.set("punctuate", "true");
// Language: if caller specified one, use it; otherwise let Deepgram auto-detect
const langParam = formData?.get("language");
if (typeof langParam === "string" && langParam.trim()) {
url.searchParams.set("language", langParam.trim());
} else {
url.searchParams.set("detect_language", "true");
}
const arrayBuffer = await file.arrayBuffer();
@@ -79,7 +135,7 @@ async function handleDeepgramTranscription(providerConfig, file, modelId, token)
method: "POST",
headers: {
...buildAuthHeaders(providerConfig, token),
"Content-Type": file.type || "audio/wav",
"Content-Type": resolveAudioContentType(file),
},
body: arrayBuffer,
});
@@ -212,7 +268,7 @@ async function handleHuggingFaceTranscription(providerConfig, file, modelId, tok
method: "POST",
headers: {
...buildAuthHeaders(providerConfig, token),
"Content-Type": file.type || "audio/wav",
"Content-Type": resolveAudioContentType(file),
},
body: arrayBuffer,
});
@@ -283,7 +339,7 @@ export async function handleAudioTranscription({
// Route to provider-specific handler
if (providerConfig.format === "deepgram") {
return handleDeepgramTranscription(providerConfig, file, modelId, token);
return handleDeepgramTranscription(providerConfig, file, modelId, token, formData);
}
if (providerConfig.format === "assemblyai") {
+28 -2
View File
@@ -308,6 +308,27 @@ export async function handleChatCore({
}
return [];
}
// (#527) tool_result → convert to text instead of dropping.
// When Claude Code + superpowers routes through Codex, it sends tool_result
// blocks in user messages. Silently dropping them causes Codex to loop
// because it never receives the tool response and keeps re-requesting it.
if (block.type === "tool_result") {
const toolId = block.tool_use_id ?? block.id ?? "unknown";
const resultContent = block.content ?? block.text ?? block.output ?? "";
const resultText =
typeof resultContent === "string"
? resultContent
: Array.isArray(resultContent)
? resultContent
.filter((c: Record<string, unknown>) => c.type === "text")
.map((c: Record<string, unknown>) => c.text)
.join("\n")
: JSON.stringify(resultContent);
if (resultText.length > 0) {
return [{ type: "text", text: `[Tool Result: ${toolId}]\n${resultText}` }];
}
return [];
}
// Unknown types: drop silently
log?.debug?.("CONTENT", `Dropped unsupported content part type="${block.type}"`);
return [];
@@ -317,10 +338,15 @@ export async function handleChatCore({
}
}
const normalizeToolCallId = getModelNormalizeToolCallId(provider || "", model || "");
const normalizeToolCallId = getModelNormalizeToolCallId(
provider || "",
model || "",
sourceFormat
);
const preserveDeveloperRole = getModelPreserveOpenAIDeveloperRole(
provider || "",
model || ""
model || "",
sourceFormat
);
translatedBody = translateRequest(
sourceFormat,
+68
View File
@@ -8,6 +8,46 @@ import {
} from "../config/constants.ts";
import { getProviderCategory } from "../config/providerRegistry.ts";
// T06 (sub2api PR #1037): Signals that indicate permanent account deactivation.
// When a 401 body contains these strings, the account is permanently dead
// and should NOT be retried after token refresh.
export const ACCOUNT_DEACTIVATED_SIGNALS = [
"account_deactivated",
"account has been deactivated",
"account has been disabled",
"your account has been suspended",
"this account is deactivated",
];
// T10 (sub2api PR #1169): Signals that indicate billing credits are exhausted.
// Distinct from rate-limit 429 — the account won't recover until credits are added.
export const CREDITS_EXHAUSTED_SIGNALS = [
"insufficient_quota",
"billing_hard_limit_reached",
"exceeded your current quota",
"credit_balance_too_low",
"your credit balance is too low",
"credits exhausted",
"out of credits",
"payment required",
];
/**
* T06: Returns true if response body indicates the account is permanently deactivated.
*/
export function isAccountDeactivated(errorText: string): boolean {
const lower = String(errorText || "").toLowerCase();
return ACCOUNT_DEACTIVATED_SIGNALS.some((sig) => lower.includes(sig));
}
/**
* T10: Returns true if response body indicates credits/quota are permanently exhausted.
*/
export function isCreditsExhausted(errorText: string): boolean {
const lower = String(errorText || "").toLowerCase();
return CREDITS_EXHAUSTED_SIGNALS.some((sig) => lower.includes(sig));
}
// ─── Provider Profile Helper ────────────────────────────────────────────────
/**
@@ -201,6 +241,14 @@ export function classifyErrorText(errorText) {
) {
return RateLimitReason.QUOTA_EXHAUSTED;
}
// T10: credits_exhausted signals
if (isCreditsExhausted(errorText)) {
return RateLimitReason.QUOTA_EXHAUSTED;
}
// T06: account_deactivated signals
if (isAccountDeactivated(errorText)) {
return RateLimitReason.AUTH_ERROR;
}
if (
lower.includes("rate limit") ||
lower.includes("too many requests") ||
@@ -301,6 +349,26 @@ export function checkFallbackError(
const errorStr = typeof errorText === "string" ? errorText : JSON.stringify(errorText);
const lowerError = errorStr.toLowerCase();
// T06 (sub2api #1037): Permanent account deactivation — do NOT retry, mark as permanent failure
if (isAccountDeactivated(errorStr)) {
return {
shouldFallback: true,
cooldownMs: 365 * 24 * 60 * 60 * 1000, // 1 year = effectively permanent
reason: RateLimitReason.AUTH_ERROR,
permanent: true,
};
}
// T10 (sub2api #1169): Credits/quota exhausted — long cooldown, distinct from rate limit
if (isCreditsExhausted(errorStr)) {
return {
shouldFallback: true,
cooldownMs: COOLDOWN_MS.paymentRequired ?? 3600 * 1000, // 1h cooldown
reason: RateLimitReason.QUOTA_EXHAUSTED,
creditsExhausted: true,
};
}
if (lowerError.includes("no credentials")) {
return {
shouldFallback: true,
+67 -4
View File
@@ -447,8 +447,10 @@ export async function handleComboChat({
const handleSingleModelWrapped = combo.context_cache_protection
? async (b, modelStr) => {
const res = await handleSingleModel(b, modelStr);
// Inject tag only on success and only for non-streaming non-binary responses
if (res.ok && !b.stream) {
if (!res.ok) return res;
// Non-streaming: inject tag into JSON response (existing logic)
if (!b.stream) {
try {
const json = await res.clone().json();
const msgs = Array.isArray(json?.messages) ? json.messages : [];
@@ -460,10 +462,71 @@ export async function handleComboChat({
});
}
} catch {
/* non-JSON or stream — skip tagging */
/* non-JSON — skip tagging */
}
return res;
}
return res;
// Streaming (Fix #490 + #511): prepend omniModel tag into the first
// non-empty content chunk so it arrives BEFORE finish_reason:stop.
// SDKs close the connection on finish_reason, so anything sent after
// that marker is silently dropped.
if (!res.body) return res;
const tagContent = `\\n<omniModel>${modelStr}</omniModel>\\n`;
const encoder = new TextEncoder();
const decoder = new TextDecoder();
let tagInjected = false;
const transform = new TransformStream({
transform(chunk, controller) {
if (tagInjected) {
// Already injected — passthrough
controller.enqueue(chunk);
return;
}
const text = decoder.decode(chunk, { stream: true });
// Look for the first SSE data line with non-empty content
// Pattern: "content":"<non-empty>" — we inject tag at the start
const contentMatch = text.match(/"content":"([^"]+)/);
if (contentMatch) {
// Inject tag at the beginning of the first content value
const injected = text.replace(
/"content":"([^"]+)/,
`"content":"${tagContent.replace(/"/g, '\\"')}$1`
);
tagInjected = true;
controller.enqueue(encoder.encode(injected));
return;
}
// No content yet — passthrough
controller.enqueue(chunk);
},
flush(controller) {
// If stream ends without ever finding content (edge case),
// inject tag as a standalone chunk before the stream closes
if (!tagInjected) {
const tagChunk = `data: ${JSON.stringify({
choices: [
{
delta: { content: tagContent },
index: 0,
finish_reason: null,
},
],
})}\n\n`;
controller.enqueue(encoder.encode(tagChunk));
}
},
});
const transformedStream = res.body.pipeThrough(transform);
return new Response(transformedStream, {
status: res.status,
headers: res.headers,
});
}
: handleSingleModel;
// ─────────────────────────────────────────────────────────────────────────
+10 -2
View File
@@ -34,7 +34,11 @@ interface Message {
// ── Context Caching Tag ─────────────────────────────────────────────────────
const CACHE_TAG_PATTERN = /<omniModel>([^<]+)<\/omniModel>/;
// Handles both actual newlines (U+000A) and literal \n sequences injected
// by combo.ts streaming around the <omniModel> tag (#531). Non-global so that
// .exec() and .test() stay stateless; callers that need full replacement use
// String.prototype.replace() which replaces all non-overlapping matches.
const CACHE_TAG_PATTERN = /(?:\\n|\n)?<omniModel>([^<]+)<\/omniModel>(?:\\n|\n)?/;
/**
* Inject the model tag into the last assistant message (or append a new one).
@@ -165,7 +169,11 @@ export function applyComboAgentMiddleware(
if (comboConfig.context_cache_protection) {
pinnedModel = extractPinnedModel(messages);
if (pinnedModel) {
// Model is pinned — caller should override model selection
// (#535) Model is pinned via <omniModel> tag — override body.model so the combo
// router uses exactly this model instead of picking a different one. Without this,
// the extracted pinnedModel is returned but body.model is unchanged, breaking
// context cache sessions by sending subsequent turns to a different model.
body = { ...body, model: pinnedModel };
}
}
+16 -2
View File
@@ -80,9 +80,17 @@ function supportsSystemRole(provider: string, model: string): boolean {
* OpenAI Responses API sends `developer`; MiniMax and most OpenAI-compatible gateways
* only accept system/user/assistant/tool and return "role param error" otherwise.
*
* Logic:
* - When targetFormat !== "openai": always convert developer system (Claude, Gemini, etc.).
* - When targetFormat === "openai": convert only when preserveDeveloperRole === false.
* This covers OpenAI-compatible providers (MiniMax, etc.) that use targetFormat "openai"
* but do not accept the developer role; the per-model preserveDeveloperRole flag is set
* via the dashboard "Compatibility" toggle ("Do not preserve developer role").
* - When targetFormat === "openai" && preserveDeveloperRole !== false: keep developer (e.g. official OpenAI).
*
* @param messages - Array of messages
* @param targetFormat - The target format (e.g., "openai", "claude", "gemini")
* @param preserveDeveloperRole - For targetFormat openai: undefined/true = keep developer (legacy default); false = map to system (MiniMax etc.)
* @param preserveDeveloperRole - For targetFormat openai: undefined/true = keep developer (legacy default); false = map to system (MiniMax and other OpenAI-compatible gateways that reject developer)
*/
export function normalizeDeveloperRole(
messages: NormalizedMessage[] | unknown,
@@ -170,8 +178,14 @@ export function normalizeSystemRole(
/**
* Full role normalization pipeline.
* Call this before sending the request to the provider.
* Applies developersystem (when needed) then systemuser for providers/models that do not support system role.
*
* @param preserveDeveloperRole - See {@link normalizeDeveloperRole}
* @param messages - Array of messages to normalize (or non-array, returned as-is)
* @param provider - Provider id for capability lookup (e.g. system role support)
* @param model - Model id for capability lookup
* @param targetFormat - Target request format (e.g. "openai", "claude", "gemini"); see {@link normalizeDeveloperRole}
* @param preserveDeveloperRole - Optional; see {@link normalizeDeveloperRole}. When false, developer role is mapped to system.
* @returns Normalized messages array, or the original value if messages is not an array
*/
export function normalizeRoles(
messages: NormalizedMessage[] | unknown,
+89
View File
@@ -173,6 +173,95 @@ export function getActiveSessions(): Array<SessionEntry & { sessionId: string; a
*/
export function clearSessions(): void {
sessions.clear();
activeSessionsByKey.clear();
}
// ─── T08: Per-API-Key Session Limit ─────────────────────────────────────────
// Tracks concurrent sticky sessions per API key and enforces max_sessions limits.
// Ref: sub2api PR #634 (fix: stabilize session hash + add user-level session limit)
// Map: apiKeyId → Set<sessionId>
const activeSessionsByKey = new Map<string, Set<string>>();
/**
* T08: Get the number of currently active sessions for an API key.
* @param apiKeyId - The API key's UUID from the database
*/
export function getActiveSessionCountForKey(apiKeyId: string): number {
return activeSessionsByKey.get(apiKeyId)?.size ?? 0;
}
/**
* T08: Register a session as belonging to an API key.
* Call this after session creation is allowed (i.e., limit check passed).
*/
export function registerKeySession(apiKeyId: string, sessionId: string): void {
if (!activeSessionsByKey.has(apiKeyId)) {
activeSessionsByKey.set(apiKeyId, new Set());
}
activeSessionsByKey.get(apiKeyId)!.add(sessionId);
}
/**
* T08: Unregister a session from an API key's active set.
* Call this when the request closes or the session TTL expires.
*/
export function unregisterKeySession(apiKeyId: string, sessionId: string): void {
activeSessionsByKey.get(apiKeyId)?.delete(sessionId);
// Clean up empty sets to avoid memory leaks
if (activeSessionsByKey.get(apiKeyId)?.size === 0) {
activeSessionsByKey.delete(apiKeyId);
}
}
/**
* T08: Check whether adding a new session would exceed the key's max_sessions limit.
* Returns null if allowed, or an error object to return as a 429 response.
*
* @param apiKeyId - The API key's UUID
* @param maxSessions - The limit from the DB (0 = unlimited)
*/
export function checkSessionLimit(
apiKeyId: string,
maxSessions: number
): { code: "SESSION_LIMIT_EXCEEDED"; message: string; limit: number; current: number } | null {
if (!maxSessions || maxSessions <= 0) return null; // unlimited
const current = getActiveSessionCountForKey(apiKeyId);
if (current < maxSessions) return null;
return {
code: "SESSION_LIMIT_EXCEEDED",
message:
`You have reached the maximum number of active sessions (${maxSessions}). ` +
`Please close unused sessions or wait for them to expire.`,
limit: maxSessions,
current,
};
}
/**
* T04: Extract an external session ID from request headers.
* Accepts both hyphenated and underscore forms for Nginx compatibility.
* Nginx drops headers with underscores by default use `underscores_in_headers on`
* in nginx.conf, or use X-Session-Id (hyphenated) which passes cleanly.
*
* Ref: sub2api README + PR #634
*
* @param headers - Request headers (Headers object or plain object with .get())
* @returns External session ID with "ext:" prefix, or null
*/
export function extractExternalSessionId(
headers: Headers | { get?: (n: string) => string | null } | null | undefined
): string | null {
if (!headers || typeof (headers as Headers).get !== "function") return null;
const h = headers as Headers;
const raw =
h.get("x-session-id") ?? // Preferred: hyphenated (passes through Nginx)
h.get("x-omniroute-session") ?? // OmniRoute-specific form
h.get("session-id") ?? // Bare session-id
null;
if (!raw || !raw.trim()) return null;
// Prefix "ext:" to ensure no collision with internal SHA-256 hash IDs
return `ext:${raw.trim().slice(0, 64)}`; // max 64 chars to avoid abuse
}
// ─── Internal Helpers ───────────────────────────────────────────────────────
+8 -2
View File
@@ -19,6 +19,8 @@ export const EFFORT_BUDGETS = {
low: 1024,
medium: 10240,
high: 131072,
max: 131072, // T11: Claude "max" / "xhigh" — full budget
xhigh: 131072, // T11: explicit alias used internally
};
// thinkingLevel string → budget token mapping
@@ -28,6 +30,8 @@ export const THINKING_LEVEL_MAP = {
low: 1024,
medium: 10240,
high: 131072,
max: 131072, // T11: max = full Claude budget (sub2api: xhigh)
xhigh: 131072, // T11: explicit xhigh alias
};
// Default config (passthrough = backward compatible)
@@ -198,7 +202,7 @@ function setCustomBudget(body, budget) {
};
}
// OpenAI reasoning_effort mapping
// OpenAI reasoning_effort mapping (T11: add 'max' tier for full budget)
if (result.reasoning_effort !== undefined || result.reasoning !== undefined) {
if (budget <= 0) {
delete result.reasoning_effort;
@@ -207,8 +211,10 @@ function setCustomBudget(body, budget) {
result.reasoning_effort = "low";
} else if (budget <= 10240) {
result.reasoning_effort = "medium";
} else {
} else if (budget < 131072) {
result.reasoning_effort = "high";
} else {
result.reasoning_effort = "max"; // T11: full budget → "max"
}
}
@@ -93,10 +93,11 @@ export function convertResponsesApiFormat(body) {
}
// Cleanup Responses API specific fields
// Note: prompt_cache_key is intentionally preserved — it is used by Codex and other
// providers as a cache-affinity signal. Stripping it breaks prompt caching (#517).
delete result.input;
delete result.instructions;
delete result.include;
delete result.prompt_cache_key;
delete result.store;
delete result.reasoning;
+7 -3
View File
@@ -96,7 +96,9 @@ export function translateRequest(
// Fix missing tool responses (insert empty tool_result if needed)
fixMissingToolResponses(result);
// Normalize roles: developer→system unless preserved, system→user for incompatible models
// Normalize roles: developer→system unless preserved, system→user for incompatible models.
// This handles (1) sourceFormat openai with messages containing developer → non-openai target
// or preserveDeveloperRole=false, and (2) all other paths where result.messages already exists.
if (result.messages && Array.isArray(result.messages)) {
result.messages = normalizeRoles(
result.messages,
@@ -151,8 +153,10 @@ export function translateRequest(
result = normalizeOpenAIResponsesRequest(result);
}
// After OPENAI_RESPONSES → OPENAI, messages are built from input; first normalizeRoles was a no-op.
// Run role pipeline again so developer→system respects preserveDeveloperRole (no hardcoding in translator).
// Second role normalization: only for OPENAI_RESPONSES. Here messages are built from input
// after the translation step, so the first normalizeRoles (above) did not see them. For
// sourceFormat openai with messages already on the body, the first block handles developer
// → system (non-openai target or preserveDeveloperRole=false); no second pass needed.
if (
sourceFormat === FORMATS.OPENAI_RESPONSES &&
result.messages &&
@@ -227,10 +227,11 @@ export function openaiResponsesToOpenAIRequest(
});
// Cleanup Responses API specific fields
// Note: prompt_cache_key is intentionally preserved — it is used by Codex and other
// providers as a cache-affinity signal. Stripping it breaks prompt caching (#517).
delete result.input;
delete result.instructions;
delete result.include;
delete result.prompt_cache_key;
delete result.store;
delete result.reasoning;
@@ -27,6 +27,60 @@ type ClaudeTool = {
defer_loading?: boolean;
};
/**
* T02: Recursively strips empty text blocks from content arrays.
* Anthropic returns 400 "text content blocks must be non-empty" if any
* text block has text: "". Must also recurse into nested tool_result.content.
* Ref: sub2api PR #1212
*/
export function stripEmptyTextBlocks(content: unknown[] | undefined): unknown[] {
if (!Array.isArray(content)) return content ?? [];
return content
.filter((block: unknown) => {
if (
block &&
typeof block === "object" &&
(block as Record<string, unknown>).type === "text"
) {
const text = (block as Record<string, unknown>).text;
if (text === "" || text == null) return false;
}
return true;
})
.map((block: unknown) => {
if (
block &&
typeof block === "object" &&
(block as Record<string, unknown>).type === "tool_result" &&
Array.isArray((block as Record<string, unknown>).content)
) {
// Recurse into nested tool_result.content
return {
...(block as Record<string, unknown>),
content: stripEmptyTextBlocks((block as Record<string, unknown>).content as unknown[]),
};
}
return block;
});
}
/**
* T15: Normalize content to string form.
* Handles both string and array-of-blocks forms (Cursor, Codex 2.x, etc.).
* Ref: sub2api PR #1197
*/
export function normalizeContentToString(content: string | unknown[] | null | undefined): string {
if (!content) return "";
if (typeof content === "string") return content;
if (Array.isArray(content)) {
return (content as Array<Record<string, unknown>>)
.filter((b) => b.type === "text")
.map((b) => String(b.text ?? ""))
.join("\n");
}
return "";
}
// Convert OpenAI request to Claude format
export function openaiToClaudeRequest(model, body, stream) {
// Check if tool prefix should be disabled (configured per-provider or global)
@@ -61,11 +115,11 @@ export function openaiToClaudeRequest(model, body, stream) {
const systemParts = [];
if (body.messages && Array.isArray(body.messages)) {
// Extract system messages
// Extract system messages (T15: handle both string and array content)
for (const msg of body.messages) {
if (msg.role === "system") {
systemParts.push(
typeof msg.content === "string" ? msg.content : extractTextContent(msg.content)
typeof msg.content === "string" ? msg.content : normalizeContentToString(msg.content)
);
}
}
@@ -270,10 +324,14 @@ function getContentBlocksFromMessage(msg, toolNameMap = new Map(), disableToolPr
const blocks = [];
if (msg.role === "tool") {
// T02: Strip empty text blocks from nested tool_result content to avoid Anthropic 400
const toolContent = Array.isArray(msg.content)
? stripEmptyTextBlocks(msg.content)
: msg.content;
blocks.push({
type: "tool_result",
tool_use_id: msg.tool_call_id,
content: msg.content,
content: toolContent,
});
} else if (msg.role === "user") {
if (typeof msg.content === "string") {
@@ -287,10 +345,14 @@ function getContentBlocksFromMessage(msg, toolNameMap = new Map(), disableToolPr
} else if (part.type === "tool_result") {
// Skip tool_result with no tool_use_id (would be useless and may cause errors)
if (!part.tool_use_id) continue;
// T02: strip empty text blocks from nested content before passing to Anthropic
const resultContent = Array.isArray(part.content)
? stripEmptyTextBlocks(part.content)
: part.content;
blocks.push({
type: "tool_result",
tool_use_id: part.tool_use_id,
content: part.content,
content: resultContent,
...(part.is_error && { is_error: part.is_error }),
});
} else if (part.type === "image_url") {
+7525 -40
View File
File diff suppressed because it is too large Load Diff
+4 -3
View File
@@ -1,6 +1,6 @@
{
"name": "omniroute",
"version": "2.8.7",
"version": "3.0.0-rc.7",
"description": "Smart AI Router with auto fallback — route to FREE & cheap models, zero downtime. Works with Cursor, Cline, Claude Desktop, Codex, and any OpenAI-compatible tool.",
"type": "module",
"bin": {
@@ -81,8 +81,10 @@
"system-info": "node scripts/system-info.mjs"
},
"dependencies": {
"@lobehub/icons": "^5.0.1",
"@modelcontextprotocol/sdk": "^1.27.1",
"@monaco-editor/react": "^4.7.0",
"@swc/helpers": "0.5.19",
"bcryptjs": "^3.0.3",
"better-sqlite3": "^12.6.2",
"bottleneck": "^2.19.5",
@@ -110,8 +112,7 @@
"uuid": "^13.0.0",
"wreq-js": "^2.0.1",
"zod": "^4.3.6",
"zustand": "^5.0.10",
"@swc/helpers": "0.5.19"
"zustand": "^5.0.10"
},
"devDependencies": {
"@playwright/test": "^1.58.2",
Binary file not shown.

After

Width:  |  Height:  |  Size: 38 KiB

+1
View File
@@ -0,0 +1 @@
<svg role="img" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg"><title>Cloudflare</title><path d="M16.5088 16.8447c.1475-.5068.0908-.9707-.1553-1.3154-.2246-.3164-.6045-.499-1.0615-.5205l-8.6592-.1123a.1559.1559 0 0 1-.1333-.0713c-.0283-.042-.0351-.0986-.021-.1553.0278-.084.1123-.1484.2036-.1562l8.7359-.1123c1.0351-.0489 2.1601-.8868 2.5537-1.9136l.499-1.3013c.0215-.0561.0293-.1128.0147-.168-.5625-2.5463-2.835-4.4453-5.5499-4.4453-2.5039 0-4.6284 1.6177-5.3876 3.8614-.4927-.3658-1.1187-.5625-1.794-.499-1.2026.119-2.1665 1.083-2.2861 2.2856-.0283.31-.0069.6128.0635.894C1.5683 13.171 0 14.7754 0 16.752c0 .1748.0142.3515.0352.5273.0141.083.0844.1475.1689.1475h15.9814c.0909 0 .1758-.0645.2032-.1553l.12-.4268zm2.7568-5.5634c-.0771 0-.1611 0-.2383.0112-.0566 0-.1054.0415-.127.0976l-.3378 1.1744c-.1475.5068-.0918.9707.1543 1.3164.2256.3164.6055.498 1.0625.5195l1.8437.1133c.0557 0 .1055.0263.1329.0703.0283.043.0351.1074.0214.1562-.0283.084-.1132.1485-.204.1553l-1.921.1123c-1.041.0488-2.1582.8867-2.5527 1.914l-.1406.3585c-.0283.0713.0215.1416.0986.1416h6.5977c.0771 0 .1474-.0489.169-.126.1122-.4082.1757-.837.1757-1.2803 0-2.6025-2.125-4.727-4.7344-4.727"/></svg>

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 14 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 18 KiB

+1
View File
@@ -0,0 +1 @@
<svg role="img" viewBox="0 0 24 24" xmlns="http://www.w3.org/2000/svg"><title>Scaleway</title><path d="M16.605 11.11v5.72a1.77 1.77 0 01-1.54 1.69h-4a1.43 1.43 0 01-1.31-1.22 1.09 1.09 0 010-.18 1.37 1.37 0 011.37-1.36h1.74a1 1 0 001-1v-3.62a1.4 1.4 0 011.18-1.39h.17a1.37 1.37 0 011.39 1.36zm-6.46 1.74V9.26a1 1 0 011-1h1.85a1.37 1.37 0 001.37-1.37 1 1 0 000-.17 1.45 1.45 0 00-1.41-1.2h-3.96a1.81 1.81 0 00-1.58 1.66v5.7a1.37 1.37 0 001.37 1.37h.21a1.4 1.4 0 001.15-1.4zm12-4.29V20a4.53 4.53 0 01-4.15 4h-7.58a8.57 8.57 0 01-8.56-8.57V4.54A4.54 4.54 0 016.395 0h7.18a8.56 8.56 0 018.56 8.56zm-2.74 0a5.83 5.83 0 00-5.82-5.82h-7.19a1.79 1.79 0 00-1.8 1.8v10.89a5.83 5.83 0 005.82 5.8h7.44a1.79 1.79 0 001.54-1.48z"/></svg>

After

Width:  |  Height:  |  Size: 723 B

+2 -1
View File
@@ -48,7 +48,8 @@ function extractChangelogSections(content) {
}
function isSemver(value) {
return /^\d+\.\d+\.\d+$/.test(value);
// Accept X.Y.Z and X.Y.Z-prerelease.N (e.g. 3.0.0-rc.1, 3.0.0-beta.2)
return /^\d+\.\d+\.\d+(-[a-zA-Z0-9.]+)?$/.test(value);
}
let hasFailure = false;
+2 -1
View File
@@ -16,7 +16,8 @@ const { dashboardPort } = runtimePorts;
const env = bootstrapEnv();
const args = ["./node_modules/next/dist/bin/next", mode, "--port", String(dashboardPort)];
if (mode === "dev") {
// Default: use webpack (stable). Set OMNIROUTE_USE_TURBOPACK=1 to use Turbopack (faster dev).
if (mode === "dev" && process.env.OMNIROUTE_USE_TURBOPACK !== "1") {
args.splice(2, 0, "--webpack");
}
@@ -0,0 +1,196 @@
/**
* Search Analytics Tab
*
* Shows search request stats from call_logs (request_type = 'search'),
* provider breakdown, cache hit rate, and cost summary.
*/
"use client";
import { useEffect, useState } from "react";
interface SearchStats {
total: number;
today: number;
cached: number;
errors: number;
totalCostUsd: number;
byProvider: Record<string, { count: number; costUsd: number }>;
last24h: Array<{ hour: string; count: number }>;
cacheHitRate: number;
avgDurationMs: number;
}
function StatCard({
icon,
label,
value,
sub,
}: {
icon: string;
label: string;
value: string | number;
sub?: string;
}) {
return (
<div className="card p-4 flex flex-col gap-1">
<div className="flex items-center gap-2 text-text-muted text-sm">
<span className="material-symbols-outlined text-[18px]">{icon}</span>
{label}
</div>
<div className="text-2xl font-bold text-text">{value}</div>
{sub && <div className="text-xs text-text-muted">{sub}</div>}
</div>
);
}
function ProviderBar({
provider,
count,
total,
costUsd,
}: {
provider: string;
count: number;
total: number;
costUsd: number;
}) {
const pct = total > 0 ? Math.round((count / total) * 100) : 0;
return (
<div className="flex flex-col gap-1">
<div className="flex justify-between text-sm">
<span className="font-medium text-text">{provider}</span>
<span className="text-text-muted">
{count} queries · ${costUsd.toFixed(4)}
</span>
</div>
<div className="h-2 rounded-full bg-bg-muted overflow-hidden">
<div
className="h-full rounded-full bg-primary transition-all"
style={{ width: `${pct}%` }}
/>
</div>
<div className="text-xs text-text-muted text-right">{pct}%</div>
</div>
);
}
export default function SearchAnalyticsTab() {
const [stats, setStats] = useState<SearchStats | null>(null);
const [loading, setLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
useEffect(() => {
fetch("/api/v1/search/analytics")
.then((r) => r.json())
.then((d) => {
setStats(d);
setLoading(false);
})
.catch((e) => {
setError(e.message);
setLoading(false);
});
}, []);
if (loading) {
return (
<div className="flex items-center justify-center py-16 text-text-muted">
<span className="material-symbols-outlined animate-spin mr-2">progress_activity</span>
Loading search analytics
</div>
);
}
if (error || !stats) {
return (
<div className="card p-6 text-center text-text-muted">
<span className="material-symbols-outlined text-[32px] mb-2 block">search_off</span>
{error || "No search data available yet."}
<p className="text-xs mt-2">
Search requests will appear here after the first search via /v1/search.
</p>
</div>
);
}
const providers = Object.entries(stats.byProvider).sort(([, a], [, b]) => b.count - a.count);
return (
<div className="flex flex-col gap-6">
{/* KPI Cards */}
<div className="grid grid-cols-2 md:grid-cols-4 gap-4">
<StatCard
icon="manage_search"
label="Total Searches"
value={stats.total.toLocaleString()}
sub={`${stats.today} today`}
/>
<StatCard
icon="cached"
label="Cache Hit Rate"
value={`${stats.cacheHitRate}%`}
sub={`${stats.cached} cached requests`}
/>
<StatCard
icon="attach_money"
label="Total Cost"
value={`$${stats.totalCostUsd.toFixed(4)}`}
sub="search API costs"
/>
<StatCard
icon="timer"
label="Avg Response"
value={`${stats.avgDurationMs}ms`}
sub={stats.errors > 0 ? `${stats.errors} errors` : "No errors"}
/>
</div>
{/* Provider Breakdown */}
{providers.length > 0 && (
<div className="card p-5">
<h3 className="font-semibold text-text mb-4 flex items-center gap-2">
<span className="material-symbols-outlined text-primary text-[20px]">hub</span>
Provider Breakdown
</h3>
<div className="flex flex-col gap-4">
{providers.map(([prov, data]) => (
<ProviderBar
key={prov}
provider={prov}
count={data.count}
total={stats.total}
costUsd={data.costUsd}
/>
))}
</div>
</div>
)}
{/* Empty state */}
{stats.total === 0 && (
<div className="card p-8 text-center text-text-muted">
<span className="material-symbols-outlined text-[48px] mb-3 block text-primary opacity-50">
travel_explore
</span>
<p className="font-medium text-text">No searches yet</p>
<p className="text-sm mt-1">
Use <code className="bg-bg-muted px-1 rounded">POST /v1/search</code> to start routing
web searches.
</p>
</div>
)}
{/* Free tier note */}
<div className="text-xs text-text-muted border border-border rounded-lg p-3 flex items-start gap-2">
<span className="material-symbols-outlined text-[16px] text-green-500 mt-0.5">
check_circle
</span>
<span>
<strong>Free tier available:</strong> Serper (2,500/mo), Brave (2,000/mo), Exa (1,000/mo),
Tavily (1,000/mo) total 6,500+ free searches/month with automatic failover.
</span>
</div>
</div>
);
}
@@ -3,15 +3,17 @@
import { useState, Suspense } from "react";
import { UsageAnalytics, CardSkeleton, SegmentedControl } from "@/shared/components";
import EvalsTab from "../usage/components/EvalsTab";
import SearchAnalyticsTab from "./SearchAnalyticsTab";
import { useTranslations } from "next-intl";
export default function AnalyticsPage() {
const [activeTab, setActiveTab] = useState("overview");
const t = useTranslations("analytics");
const tabDescriptions = {
const tabDescriptions: Record<string, string> = {
overview: t("overviewDescription"),
evals: t("evalsDescription"),
search: "Search request analytics — provider breakdown, cache hit rate, and cost tracking.",
};
return (
@@ -29,6 +31,7 @@ export default function AnalyticsPage() {
options={[
{ value: "overview", label: t("overview") },
{ value: "evals", label: t("evals") },
{ value: "search", label: "Search" },
]}
value={activeTab}
onChange={setActiveTab}
@@ -40,6 +43,7 @@ export default function AnalyticsPage() {
</Suspense>
)}
{activeTab === "evals" && <EvalsTab />}
{activeTab === "search" && <SearchAnalyticsTab />}
</div>
);
}
@@ -523,15 +523,12 @@ export default function ApiManagerPageClient() {
</div>
<div className="col-span-3 flex items-center gap-1.5">
<code className="text-sm text-text-muted font-mono truncate">{key.key}</code>
<button
onClick={() => copy(key.key, key.id)}
className="p-1 hover:bg-black/5 dark:hover:bg-white/5 rounded text-text-muted hover:text-primary opacity-0 group-hover:opacity-100 transition-all shrink-0"
title={t("copyMaskedKey")}
<span
className="p-1 text-text-muted/40 opacity-0 group-hover:opacity-100 transition-all shrink-0 cursor-help"
title={t("keyOnlyAvailableAtCreation")}
>
<span className="material-symbols-outlined text-[14px]">
{copied === key.id ? "check" : "content_copy"}
</span>
</button>
<span className="material-symbols-outlined text-[14px]">lock</span>
</span>
</div>
<div className="col-span-2 flex items-center">
<div className="flex flex-col items-start gap-1">
@@ -355,25 +355,33 @@ export default function AntigravityToolCard({
)}
{/* When stopped: how it works */}
{!isRunning && (
<div className="flex flex-col gap-1.5 px-1">
<p className="text-xs text-text-muted">
<span className="font-medium text-text-main">{t("howItWorks")}</span>{" "}
{t("antigravityHowWorksDesc")}
</p>
<div className="flex flex-col gap-0.5 text-[11px] text-text-muted">
<span>{t("antigravityStep1")}</span>
<span>
{t("antigravityStep2Prefix")}{" "}
<code className="text-[10px] bg-surface px-1 rounded">
daily-cloudcode-pa.googleapis.com
</code>{" "}
{t("antigravityStep2Suffix")}
</span>
<span>{t("antigravityStep3")}</span>
</div>
</div>
)}
{!isRunning &&
(() => {
// Dynamic MITM instructions per tool (#505)
const mitmDomains: Record<string, string> = {
antigravity: "daily-cloudcode-pa.googleapis.com",
kiro: "api.anthropic.com",
};
const toolName = tool.name || tool.id;
const domain = mitmDomains[tool.id] || mitmDomains.antigravity;
return (
<div className="flex flex-col gap-1.5 px-1">
<p className="text-xs text-text-muted">
<span className="font-medium text-text-main">{t("howItWorks")}</span>{" "}
{t("mitmHowWorksDesc", { toolName })}
</p>
<div className="flex flex-col gap-0.5 text-[11px] text-text-muted">
<span>{t("mitmStep1")}</span>
<span>
{t("mitmStep2Prefix")}{" "}
<code className="text-[10px] bg-surface px-1 rounded">{domain}</code>{" "}
{t("mitmStep2Suffix")}
</span>
<span>{t("mitmStep3", { toolName })}</span>
</div>
</div>
);
})()}
</div>
)}
@@ -58,8 +58,10 @@ export default function ClaudeToolCard({
const effectiveConfigStatus = configStatus || batchStatus?.configStatus || null;
useEffect(() => {
// (#523) Store the key *id* (not the masked string) so the backend can
// resolve the real secret from DB before writing to settings.json.
if (apiKeys?.length > 0 && !selectedApiKey) {
setSelectedApiKey(apiKeys[0].key);
setSelectedApiKey(apiKeys[0].id);
}
}, [apiKeys, selectedApiKey]);
@@ -95,10 +97,11 @@ export default function ClaudeToolCard({
}
}
});
// Only set selectedApiKey if it exists in apiKeys list
// Restore selected key from file: match token stored in file against known keys
const tokenFromFile = env.ANTHROPIC_AUTH_TOKEN;
if (tokenFromFile && apiKeys?.some((k) => k.key === tokenFromFile)) {
setSelectedApiKey(tokenFromFile);
if (tokenFromFile) {
const matchedKey = apiKeys?.find((k) => k.key === tokenFromFile);
if (matchedKey) setSelectedApiKey(matchedKey.id);
}
}
}, [claudeStatus, apiKeys, tool.defaultModels, onModelMappingChange]);
@@ -132,24 +135,27 @@ export default function ClaudeToolCard({
try {
const env: any = { ANTHROPIC_BASE_URL: getEffectiveBaseUrl() };
// Get key from dropdown, fallback to first key or sk_omniroute for localhost
const keyToUse =
selectedApiKey?.trim() ||
(apiKeys?.length > 0 ? apiKeys[0].key : null) ||
(!cloudEnabled ? "sk_omniroute" : null);
// (#523) Prefer keyId lookup so the backend writes the real key to disk.
// Fall back to sk_omniroute for localhost-only setups without a key.
const selectedKeyId = selectedApiKey?.trim() || (apiKeys?.length > 0 ? apiKeys[0].id : null);
const skOmnirouteFallback = !cloudEnabled ? "sk_omniroute" : null;
if (keyToUse) {
env.ANTHROPIC_AUTH_TOKEN = keyToUse;
if (!selectedKeyId && skOmnirouteFallback) {
env.ANTHROPIC_AUTH_TOKEN = skOmnirouteFallback;
}
tool.defaultModels.forEach((model) => {
const targetModel = modelMappings[model.alias];
if (targetModel && model.envKey) env[model.envKey] = targetModel;
});
const postBody: Record<string, unknown> = { env };
if (selectedKeyId) postBody.keyId = selectedKeyId;
const res = await fetch("/api/cli-tools/claude-settings", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ env }),
body: JSON.stringify(postBody),
});
const data = await res.json();
if (res.ok) {
@@ -412,7 +418,7 @@ export default function ClaudeToolCard({
className="flex-1 px-2 py-1.5 bg-surface rounded text-xs border border-border focus:outline-none focus:ring-1 focus:ring-primary/50"
>
{apiKeys.map((key) => (
<option key={key.id} value={key.key}>
<option key={key.id} value={key.id}>
{key.key}
</option>
))}
File diff suppressed because it is too large Load Diff
@@ -2,6 +2,7 @@
import { useState, useEffect } from "react";
import Image from "next/image";
import ProviderIcon from "@/shared/components/ProviderIcon";
import PropTypes from "prop-types";
import {
Card,
@@ -490,16 +491,8 @@ function ProviderCard({ providerId, provider, stats, authType, onToggle }) {
const t = useTranslations("providers");
const tc = useTranslations("common");
const { connected, error, errorCode, errorTime, allDisabled } = stats;
const [imgSrc, setImgSrc] = useState(`/providers/${provider.id}.png`);
const [imgError, setImgError] = useState(false);
const handleImgError = () => {
if (imgSrc.endsWith(".png")) {
setImgSrc(`/providers/${provider.id}.svg`);
} else {
setImgError(true);
}
};
// (#529) Icon state replaced by ProviderIcon component (Lobehub + PNG + generic fallback)
const dotColors = {
free: "bg-green-500",
@@ -526,21 +519,8 @@ function ProviderCard({ providerId, provider, stats, authType, onToggle }) {
className="size-8 rounded-lg flex items-center justify-center"
style={{ backgroundColor: `${provider.color}15` }}
>
{imgError ? (
<span className="text-xs font-bold" style={{ color: provider.color }}>
{provider.textIcon || provider.id.slice(0, 2).toUpperCase()}
</span>
) : (
<Image
src={imgSrc}
alt={provider.name}
width={30}
height={30}
className="object-contain rounded-lg max-w-[32px] max-h-[32px]"
sizes="32px"
onError={handleImgError}
/>
)}
{/* (#529) ProviderIcon: Lobehub icons → PNG fallback → generic icon */}
<ProviderIcon providerId={provider.id} size={28} type="color" />
</div>
<div>
<h3 className="font-semibold flex items-center gap-1.5">
@@ -633,28 +613,15 @@ function ApiKeyProviderCard({ providerId, provider, stats, authType, onToggle })
compatible: t("compatibleLabel"),
};
// Determine icon path: OpenAI Compatible providers use specialized icons
const getIconPath = () => {
// (#529) Icon state replaced by ProviderIcon component
// For compatible/anthropic providers, continue using static PNGs via the icon path
const staticIconPath = (() => {
if (isCompatible) {
return provider.apiType === "responses" ? "/providers/oai-r.png" : "/providers/oai-cc.png";
}
if (isAnthropicCompatible) {
return "/providers/anthropic-m.png"; // Use Anthropic icon as base
}
return `/providers/${provider.id}.png`;
};
const [imgSrc, setImgSrc] = useState<string>(() => getIconPath());
const [imgError, setImgError] = useState(false);
const handleImgError = () => {
const basePath = getIconPath();
if (imgSrc.endsWith(".png") && !isCompatible && !isAnthropicCompatible) {
setImgSrc(`/providers/${provider.id}.svg`);
} else {
setImgError(true);
}
};
if (isAnthropicCompatible) return "/providers/anthropic-m.png";
return null; // ProviderIcon will handle it
})();
return (
<Link href={`/dashboard/providers/${providerId}`} className="group">
@@ -668,20 +635,18 @@ function ApiKeyProviderCard({ providerId, provider, stats, authType, onToggle })
className="size-8 rounded-lg flex items-center justify-center"
style={{ backgroundColor: `${provider.color}15` }}
>
{imgError ? (
<span className="text-xs font-bold" style={{ color: provider.color }}>
{provider.textIcon || provider.id.slice(0, 2).toUpperCase()}
</span>
) : (
{/* (#529) ProviderIcon with static override for compatible providers */}
{staticIconPath ? (
<Image
src={imgSrc || getIconPath()}
src={staticIconPath}
alt={provider.name}
width={30}
height={30}
className="object-contain rounded-lg max-w-[30px] max-h-[30px]"
sizes="30px"
onError={handleImgError}
/>
) : (
<ProviderIcon providerId={provider.id} size={28} type="color" />
)}
</div>
<div>
@@ -12,6 +12,7 @@ import { createBackup } from "@/shared/services/backupService";
import { saveCliToolLastConfigured, deleteCliToolLastConfigured } from "@/lib/db/cliToolState";
import { cliSettingsEnvSchema } from "@/shared/validation/schemas";
import { isValidationFailure, validateBody } from "@/shared/validation/helpers";
import { getApiKeyById } from "@/lib/localDb";
// Get claude settings path based on OS
const getClaudeSettingsPath = () => getCliPrimaryConfigPath("claude");
@@ -100,6 +101,22 @@ export async function POST(request: Request) {
}
const { env } = validation.data;
// (#523/#526) If a keyId was provided, resolve the real API key from DB.
// The /api/keys list endpoint returns masked key strings — sending those to
// disk would save an unusable half-hidden token. Resolving by ID guarantees
// we always write the full key value to the config file.
const keyId = typeof rawBody?.keyId === "string" ? rawBody.keyId.trim() : null;
if (keyId) {
try {
const keyRecord = await getApiKeyById(keyId);
if (keyRecord?.key) {
env.ANTHROPIC_AUTH_TOKEN = keyRecord.key as string;
}
} catch {
// Non-critical: fall back to whatever value was in env (e.g. sk_omniroute)
}
}
const settingsPath = getClaudeSettingsPath();
const claudeDir = path.dirname(settingsPath);
+13 -1
View File
@@ -9,6 +9,7 @@ import { createBackup } from "@/shared/services/backupService";
import { saveCliToolLastConfigured, deleteCliToolLastConfigured } from "@/lib/db/cliToolState";
import { cliModelConfigSchema } from "@/shared/validation/schemas";
import { isValidationFailure, validateBody } from "@/shared/validation/helpers";
import { getApiKeyById } from "@/lib/localDb";
const CLINE_DATA_DIR = path.join(os.homedir(), ".cline", "data");
const GLOBAL_STATE_PATH = path.join(CLINE_DATA_DIR, "globalState.json");
@@ -125,7 +126,18 @@ export async function POST(request: Request) {
if (isValidationFailure(validation)) {
return NextResponse.json({ error: validation.error }, { status: 400 });
}
const { baseUrl, apiKey, model } = validation.data;
let { baseUrl, apiKey, model } = validation.data;
// (#526) Resolve real key from DB if keyId was provided
const keyId = typeof rawBody?.keyId === "string" ? rawBody.keyId.trim() : null;
if (keyId) {
try {
const keyRecord = await getApiKeyById(keyId);
if (keyRecord?.key) apiKey = keyRecord.key as string;
} catch {
/* non-critical */
}
}
// Ensure directory exists
await fs.mkdir(CLINE_DATA_DIR, { recursive: true });
@@ -39,6 +39,10 @@ export async function POST(request, { params }) {
switch (toolId) {
case "continue":
return await saveContinueConfig({ baseUrl, apiKey, model });
case "opencode":
// (#524) OpenCode config was never saved because only 'continue' was handled here.
// opencode reads ~/.config/opencode/config.toml — write the OmniRoute settings there.
return await saveOpenCodeConfig({ baseUrl, apiKey, model });
default:
return NextResponse.json(
{ error: `Direct config save not supported for: ${toolId}` },
@@ -125,3 +129,56 @@ async function saveContinueConfig({ baseUrl, apiKey, model }) {
configPath,
});
}
/**
* Save OpenCode config to ~/.config/opencode/config.toml (XDG_CONFIG_HOME aware).
* (#524) OpenCode was silently failing because this handler was missing.
*/
async function saveOpenCodeConfig({ baseUrl, apiKey, model }) {
const { apiPort } = getRuntimePorts();
// Honour $XDG_CONFIG_HOME if set, otherwise use ~/.config per the XDG Base Directory spec
const xdgConfigHome = process.env.XDG_CONFIG_HOME || path.join(os.homedir(), ".config");
const configPath = path.join(xdgConfigHome, "opencode", "config.toml");
const configDir = path.dirname(configPath);
// Ensure ~/.config/opencode/ exists
await fs.mkdir(configDir, { recursive: true });
const normalizedBaseUrl = String(baseUrl || "")
.trim()
.replace(/\/+$/, "");
// Read existing TOML to preserve any user settings outside our block
let existingContent = "";
try {
existingContent = await fs.readFile(configPath, "utf-8");
} catch {
// File doesn't exist yet — start fresh
}
// Build the OmniRoute TOML block.
// opencode config.toml uses the [provider.X] table format.
void apiPort; // available for future port-based detection
const omniBlock = `
# OmniRoute managed updated automatically by OmniRoute CLI Tools
[provider.omniroute]
api_key = "${apiKey || "sk_omniroute"}"
base_url = "${normalizedBaseUrl}"
model = "${model}"
`;
// Remove old OmniRoute-managed block (if any) then append fresh one
const cleanedContent = existingContent
.replace(/\n?# OmniRoute managed[\s\S]*?(?=\n\[|$)/, "")
.trimEnd();
const newContent = (cleanedContent ? cleanedContent + "\n" : "") + omniBlock;
await fs.writeFile(configPath, newContent, "utf-8");
return NextResponse.json({
success: true,
message: `OpenCode config saved to ${configPath}`,
configPath,
});
}
@@ -12,6 +12,7 @@ import { createBackup } from "@/shared/services/backupService";
import { saveCliToolLastConfigured, deleteCliToolLastConfigured } from "@/lib/db/cliToolState";
import { cliModelConfigSchema } from "@/shared/validation/schemas";
import { isValidationFailure, validateBody } from "@/shared/validation/helpers";
import { getApiKeyById } from "@/lib/localDb";
const getOpenClawSettingsPath = () => getCliPrimaryConfigPath("openclaw");
const getOpenClawDir = () => path.dirname(getOpenClawSettingsPath());
@@ -101,7 +102,18 @@ export async function POST(request: Request) {
if (isValidationFailure(validation)) {
return NextResponse.json({ error: validation.error }, { status: 400 });
}
const { baseUrl, apiKey, model } = validation.data;
let { baseUrl, apiKey, model } = validation.data;
// (#526) Resolve real key from DB if keyId was provided
const keyId = typeof rawBody?.keyId === "string" ? rawBody.keyId.trim() : null;
if (keyId) {
try {
const keyRecord = await getApiKeyById(keyId);
if (keyRecord?.key) apiKey = keyRecord.key as string;
} catch {
/* non-critical */
}
}
const openclawDir = getOpenClawDir();
const settingsPath = getOpenClawSettingsPath();
+48 -14
View File
@@ -6,7 +6,13 @@ import {
updateCustomModel,
getModelCompatOverrides,
mergeModelCompatOverride,
type ModelCompatPatch,
} from "@/lib/localDb";
import {
AI_PROVIDERS,
isOpenAICompatibleProvider,
isAnthropicCompatibleProvider,
} from "@/shared/constants/providers";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { providerModelMutationSchema } from "@/shared/validation/schemas";
import { isValidationFailure, validateBody } from "@/shared/validation/helpers";
@@ -32,9 +38,9 @@ export async function GET(request) {
const modelCompatOverrides = provider ? getModelCompatOverrides(provider) : [];
return Response.json({ models, modelCompatOverrides });
} catch (error) {
} catch {
return Response.json(
{ error: { message: error.message, type: "server_error" } },
{ error: { message: "Failed to fetch provider models", type: "server_error" } },
{ status: 500 }
);
}
@@ -124,6 +130,7 @@ export async function PUT(request) {
supportedEndpoints,
normalizeToolCallId,
preserveOpenAIDeveloperRole,
compatByProtocol,
} = validation.data;
const raw = rawBody as Record<string, unknown>;
@@ -134,6 +141,9 @@ export async function PUT(request) {
if ("normalizeToolCallId" in raw) updates.normalizeToolCallId = normalizeToolCallId;
if ("preserveOpenAIDeveloperRole" in raw)
updates.preserveOpenAIDeveloperRole = preserveOpenAIDeveloperRole;
if ("compatByProtocol" in raw && compatByProtocol !== undefined) {
updates.compatByProtocol = compatByProtocol;
}
const model = await updateCustomModel(provider, modelId, updates);
@@ -142,24 +152,48 @@ export async function PUT(request) {
const compatOnly =
rawKeys.length > 0 &&
rawKeys.every((k) =>
["provider", "modelId", "normalizeToolCallId", "preserveOpenAIDeveloperRole"].includes(k)
[
"provider",
"modelId",
"normalizeToolCallId",
"preserveOpenAIDeveloperRole",
"compatByProtocol",
].includes(k)
) &&
("normalizeToolCallId" in raw || "preserveOpenAIDeveloperRole" in raw);
("normalizeToolCallId" in raw ||
"preserveOpenAIDeveloperRole" in raw ||
"compatByProtocol" in raw);
if (compatOnly) {
const patch: {
normalizeToolCallId?: boolean;
preserveOpenAIDeveloperRole?: boolean;
} = {};
const knownProvider =
!!provider &&
(Object.prototype.hasOwnProperty.call(
AI_PROVIDERS as Record<string, unknown>,
provider
) ||
isOpenAICompatibleProvider(provider) ||
isAnthropicCompatibleProvider(provider));
if (!knownProvider) {
return Response.json(
{ error: { message: "Unknown provider", type: "validation_error" } },
{ status: 400 }
);
}
const patch: ModelCompatPatch = {};
if ("normalizeToolCallId" in raw && typeof normalizeToolCallId === "boolean") {
patch.normalizeToolCallId = normalizeToolCallId;
}
if (
"preserveOpenAIDeveloperRole" in raw &&
typeof preserveOpenAIDeveloperRole === "boolean"
) {
patch.preserveOpenAIDeveloperRole = preserveOpenAIDeveloperRole;
if ("preserveOpenAIDeveloperRole" in raw) {
patch.preserveOpenAIDeveloperRole =
preserveOpenAIDeveloperRole === null || typeof preserveOpenAIDeveloperRole === "boolean"
? preserveOpenAIDeveloperRole
: undefined;
}
if ("compatByProtocol" in raw && compatByProtocol && typeof compatByProtocol === "object") {
patch.compatByProtocol = compatByProtocol;
}
if (Object.keys(patch).length > 0) {
mergeModelCompatOverride(provider, modelId, patch);
}
mergeModelCompatOverride(provider, modelId, patch);
return Response.json({
ok: true,
modelCompatOverrides: getModelCompatOverrides(provider),
+48 -8
View File
@@ -90,13 +90,23 @@ export async function POST(request) {
});
}
// Test each connection sequentially via direct function call (no HTTP self-call)
const results = [];
// Test each connection with timeout and concurrency limits (prevents server crash on large groups)
const PER_CONNECTION_TIMEOUT = 30_000; // 30s per connection
const CONCURRENCY = 5; // max parallel tests
for (const conn of connectionsToTest) {
const testOne = async (conn) => {
try {
const data = await testSingleConnection(conn.id);
results.push({
const result = await Promise.race([
testSingleConnection(conn.id),
new Promise((_, reject) =>
setTimeout(
() => reject(new Error("Connection test timed out after 30s")),
PER_CONNECTION_TIMEOUT
)
),
]);
const data = result as any;
return {
provider: conn.provider,
connectionId: conn.id,
connectionName: conn.name || conn.email || conn.provider,
@@ -107,9 +117,9 @@ export async function POST(request) {
diagnosis: data.diagnosis || null,
statusCode: data.statusCode || null,
testedAt: data.testedAt || new Date().toISOString(),
});
};
} catch (error) {
results.push({
return {
provider: conn.provider,
connectionId: conn.id,
connectionName: conn.name || conn.email || conn.provider,
@@ -120,7 +130,37 @@ export async function POST(request) {
diagnosis: { type: "network_error", source: "local", code: null, message: error.message },
statusCode: null,
testedAt: new Date().toISOString(),
});
};
}
};
// Execute with concurrency limit
const results = [];
for (let i = 0; i < connectionsToTest.length; i += CONCURRENCY) {
const batch = connectionsToTest.slice(i, i + CONCURRENCY);
const batchResults = await Promise.allSettled(batch.map(testOne));
for (const r of batchResults) {
results.push(
r.status === "fulfilled"
? r.value
: {
provider: "unknown",
connectionId: "unknown",
connectionName: "unknown",
authType: "unknown",
valid: false,
latencyMs: 0,
error: r.reason?.message || "Test failed",
diagnosis: {
type: "network_error",
source: "local",
code: null,
message: r.reason?.message || "Test failed",
},
statusCode: null,
testedAt: new Date().toISOString(),
}
);
}
}
+11
View File
@@ -1,7 +1,9 @@
import { NextResponse } from "next/server";
import initializeCloudSync from "@/shared/services/initializeCloudSync";
import { startModelSyncScheduler } from "@/shared/services/modelSyncScheduler";
let syncInitialized = false;
let modelSyncInitialized = false;
// POST /api/sync/initialize - Initialize cloud sync scheduler
export async function POST(request) {
@@ -15,9 +17,17 @@ export async function POST(request) {
await initializeCloudSync();
syncInitialized = true;
// (#488) Start model auto-sync scheduler (24h, configurable via MODEL_SYNC_INTERVAL_HOURS)
if (!modelSyncInitialized) {
const origin = request.headers.get("origin") || "http://localhost:20128";
startModelSyncScheduler(origin);
modelSyncInitialized = true;
}
return NextResponse.json({
success: true,
message: "Cloud sync initialized successfully",
modelSyncEnabled: true,
});
} catch (error) {
console.log("Error initializing cloud sync:", error);
@@ -34,6 +44,7 @@ export async function POST(request) {
export async function GET(request) {
return NextResponse.json({
initialized: syncInitialized,
modelSyncInitialized,
message: syncInitialized ? "Cloud sync is running" : "Cloud sync not initialized",
});
}
+37 -3
View File
@@ -1,4 +1,8 @@
import { getProviderConnectionById, updateProviderConnection, resolveProxyForConnection } from "@/lib/localDb";
import {
getProviderConnectionById,
updateProviderConnection,
resolveProxyForConnection,
} from "@/lib/localDb";
import { getMachineId } from "@/shared/utils/machine";
import { getUsageForProvider } from "@omniroute/open-sse/services/usage.ts";
import { getExecutor } from "@omniroute/open-sse/executors/index.ts";
@@ -109,7 +113,10 @@ async function refreshAndUpdateCredentials(connection: any) {
/**
* GET /api/usage/[connectionId] - Get usage data for a specific connection
*/
export async function GET(request: Request, { params }: { params: Promise<{ connectionId: string }> }) {
export async function GET(
request: Request,
{ params }: { params: Promise<{ connectionId: string }> }
) {
try {
const { connectionId } = await params;
@@ -155,7 +162,34 @@ export async function GET(request: Request, { params }: { params: Promise<{ conn
// Populate quota cache for quota-aware account selection
if (isRecord(usage?.quotas)) {
setQuotaCache(connectionId, connection.provider, usage.quotas);
setQuotaCache(
connectionId,
connection.provider as string,
usage.quotas as Record<string, unknown>
);
}
// (#491) If the live usage check returned an auth error, sync the expired status
// back to the DB so the Providers page reflects the same degraded state as
// Limits & Quotas (which performs the live check).
const errorMessage = typeof usage?.message === "string" ? usage.message.toLowerCase() : "";
const isAuthError =
errorMessage.includes("token expired") ||
errorMessage.includes("access denied") ||
errorMessage.includes("re-authenticate") ||
errorMessage.includes("unauthorized");
if (isAuthError && connection.testStatus !== "expired") {
try {
await updateProviderConnection(connection.id as string, {
testStatus: "expired",
lastErrorType: "token_expired",
lastErrorAt: new Date().toISOString(),
});
} catch (dbErr) {
// Non-critical: log but don't block the response
console.error("[Usage API] Failed to sync expired status to DB:", dbErr);
}
}
return Response.json(usage);
@@ -0,0 +1,49 @@
import { NextResponse } from "next/server";
import { z } from "zod";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { setAccountKeyLimit, getAccountKeyLimit } from "@/lib/db/registeredKeys";
const limitsSchema = z.object({
maxActiveKeys: z.number().int().positive().nullable().optional(),
dailyIssueLimit: z.number().int().positive().nullable().optional(),
hourlyIssueLimit: z.number().int().positive().nullable().optional(),
});
/**
* GET /api/v1/accounts/[id]/limits
* Get the current issuance limits for an account.
*/
export async function GET(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const limits = getAccountKeyLimit(params.id);
return NextResponse.json({ accountId: params.id, limits: limits ?? null });
}
/**
* PUT /api/v1/accounts/[id]/limits
* Configure issuance limits for an account.
*/
export async function PUT(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
let body: unknown;
try {
body = await request.json();
} catch {
return NextResponse.json({ error: "Invalid JSON body" }, { status: 400 });
}
const parsed = limitsSchema.safeParse(body);
if (!parsed.success) {
return NextResponse.json({ error: parsed.error.flatten() }, { status: 400 });
}
setAccountKeyLimit(params.id, parsed.data);
const updated = getAccountKeyLimit(params.id);
return NextResponse.json({ accountId: params.id, limits: updated });
}
+3 -1
View File
@@ -160,6 +160,7 @@ export async function POST(request) {
// Resolve provider config — dynamic first (local override), then hardcoded
let providerConfig: EmbeddingProvider | null =
dynamicProviders.find((dp) => dp.id === provider) || getEmbeddingProvider(provider) || null;
let credentialsProviderId = provider;
// #496: Fallback — resolve from ALL provider_nodes (not just localhost)
// This enables custom embedding models (e.g. google/gemini-embedding-001) whose
@@ -180,6 +181,7 @@ export async function POST(request) {
authHeader: "bearer",
models: [],
};
credentialsProviderId = matchingNode.id || provider;
log.info(
"EMBED",
`Resolved custom embedding provider: ${provider}${providerConfig.baseUrl}`
@@ -200,7 +202,7 @@ export async function POST(request) {
// Get credentials — skip for local providers (authType: "none")
let credentials = null;
if (providerConfig && providerConfig.authType !== "none") {
credentials = await getProviderCredentials(provider);
credentials = await getProviderCredentials(credentialsProviderId);
if (!credentials) {
return errorResponse(
HTTP_STATUS.BAD_REQUEST,
+123
View File
@@ -0,0 +1,123 @@
import { NextResponse } from "next/server";
import { z } from "zod";
import { isAuthenticated } from "@/shared/utils/apiAuth";
const reportSchema = z.object({
title: z.string().min(1).max(300),
provider: z.string().max(80).optional(),
accountId: z.string().max(120).optional(),
requestId: z.string().max(200).optional(),
errorCode: z.string().max(100).optional(),
details: z.record(z.unknown()).optional(),
labels: z.array(z.string().max(50)).optional(),
});
/**
* POST /api/v1/issues/report
*
* Optionally report a quota-exceeded or key-issuance failure event to GitHub.
*
* Requires GITHUB_ISSUES_REPO (format: owner/repo) and GITHUB_ISSUES_TOKEN
* environment variables to be set. If not configured, returns 202 (accepted but
* logged only).
*/
export async function POST(request: Request) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
let body: unknown;
try {
body = await request.json();
} catch {
return NextResponse.json({ error: "Invalid JSON body" }, { status: 400 });
}
const parsed = reportSchema.safeParse(body);
if (!parsed.success) {
return NextResponse.json({ error: parsed.error.flatten() }, { status: 400 });
}
const { title, provider, accountId, requestId, errorCode, details, labels = [] } = parsed.data;
const repo = process.env.GITHUB_ISSUES_REPO;
const token = process.env.GITHUB_ISSUES_TOKEN;
// ── Structured body for the GitHub issue ──
const issueBody = [
`## ${errorCode ?? "Key Issuance Event"}`,
"",
"| Field | Value |",
"|-------|-------|",
provider ? `| Provider | \`${provider}\` |` : null,
accountId ? `| Account ID | \`${accountId}\` |` : null,
requestId ? `| Request ID | \`${requestId}\` |` : null,
errorCode ? `| Error Code | \`${errorCode}\` |` : null,
`| Reported At | ${new Date().toISOString()} |`,
"",
details ? "### Details\n```json\n" + JSON.stringify(details, null, 2) + "\n```" : null,
"",
"_Auto-reported by OmniRoute Registered Key Issuer_",
]
.filter(Boolean)
.join("\n");
// ── Log locally regardless ──
console.log(
`[issues/report] title="${title}" errorCode=${errorCode ?? "—"} provider=${provider ?? "—"} accountId=${accountId ?? "—"}`
);
if (!repo || !token) {
// No GitHub config — log only
return NextResponse.json(
{
logged: true,
githubIssueCreated: false,
reason: !repo ? "GITHUB_ISSUES_REPO not configured" : "GITHUB_ISSUES_TOKEN not configured",
},
{ status: 202 }
);
}
// ── Create GitHub issue ──
try {
const [owner, repoName] = repo.split("/");
const ghRes = await fetch(`https://api.github.com/repos/${owner}/${repoName}/issues`, {
method: "POST",
headers: {
Authorization: `Bearer ${token}`,
Accept: "application/vnd.github+json",
"X-GitHub-Api-Version": "2022-11-28",
"Content-Type": "application/json",
},
body: JSON.stringify({
title: `[Key Issuer] ${title}`,
body: issueBody,
labels: ["key-issuer", "automated", ...labels],
}),
});
if (!ghRes.ok) {
const errText = await ghRes.text();
console.error(`[issues/report] GitHub API error ${ghRes.status}: ${errText}`);
return NextResponse.json(
{ logged: true, githubIssueCreated: false, githubError: ghRes.status },
{ status: 207 }
);
}
const ghData = await ghRes.json();
return NextResponse.json({
logged: true,
githubIssueCreated: true,
githubIssueUrl: ghData.html_url,
githubIssueNumber: ghData.number,
});
} catch (err) {
console.error("[issues/report] GitHub fetch failed:", err);
return NextResponse.json(
{ logged: true, githubIssueCreated: false, error: "GitHub request failed" },
{ status: 207 }
);
}
}
@@ -0,0 +1,49 @@
import { NextResponse } from "next/server";
import { z } from "zod";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { setProviderKeyLimit, getProviderKeyLimit } from "@/lib/db/registeredKeys";
const limitsSchema = z.object({
maxActiveKeys: z.number().int().positive().nullable().optional(),
dailyIssueLimit: z.number().int().positive().nullable().optional(),
hourlyIssueLimit: z.number().int().positive().nullable().optional(),
});
/**
* GET /api/v1/providers/[id]/limits
* Get the current issuance limits for a provider.
*/
export async function GET(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const limits = getProviderKeyLimit(params.id);
return NextResponse.json({ provider: params.id, limits: limits ?? null });
}
/**
* PUT /api/v1/providers/[id]/limits
* Configure issuance limits for a provider.
*/
export async function PUT(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
let body: unknown;
try {
body = await request.json();
} catch {
return NextResponse.json({ error: "Invalid JSON body" }, { status: 400 });
}
const parsed = limitsSchema.safeParse(body);
if (!parsed.success) {
return NextResponse.json({ error: parsed.error.flatten() }, { status: 400 });
}
setProviderKeyLimit(params.id, parsed.data);
const updated = getProviderKeyLimit(params.id);
return NextResponse.json({ provider: params.id, limits: updated });
}
+33
View File
@@ -0,0 +1,33 @@
import { NextResponse } from "next/server";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { checkQuota } from "@/lib/db/registeredKeys";
/**
* GET /api/v1/quotas/check?provider=&accountId=
*
* Check if a new registered key can be issued for the given provider/account
* without actually issuing one. Use this to pre-validate before POST /registered-keys.
*/
export async function GET(request: Request) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const { searchParams } = new URL(request.url);
const provider = searchParams.get("provider") ?? "";
const accountId = searchParams.get("accountId") ?? "";
try {
const result = checkQuota(provider, accountId);
return NextResponse.json({
allowed: result.allowed,
...(result.errorCode ? { errorCode: result.errorCode, reason: result.errorMessage } : {}),
provider: provider || null,
accountId: accountId || null,
checkedAt: new Date().toISOString(),
});
} catch (err) {
console.error("[quotas/check] error:", err);
return NextResponse.json({ error: "Quota check failed" }, { status: 500 });
}
}
@@ -0,0 +1,21 @@
import { NextResponse } from "next/server";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { revokeRegisteredKey } from "@/lib/db/registeredKeys";
/**
* POST /api/v1/registered-keys/[id]/revoke
*
* Explicit revoke endpoint (supports clients that cannot issue DELETE requests).
*/
export async function POST(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const revoked = revokeRegisteredKey(params.id);
if (!revoked) {
return NextResponse.json({ error: "Key not found or already revoked" }, { status: 404 });
}
return NextResponse.json({ success: true, id: params.id, revokedAt: new Date().toISOString() });
}
@@ -0,0 +1,33 @@
import { NextResponse } from "next/server";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { getRegisteredKey, revokeRegisteredKey } from "@/lib/db/registeredKeys";
// ─── GET /api/v1/registered-keys/[id] ────────────────────────────────────────
export async function GET(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const key = getRegisteredKey(params.id);
if (!key) {
return NextResponse.json({ error: "Key not found" }, { status: 404 });
}
return NextResponse.json({ key });
}
// ─── DELETE /api/v1/registered-keys/[id] ─────────────────────────────────────
export async function DELETE(request: Request, { params }: { params: { id: string } }) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const revoked = revokeRegisteredKey(params.id);
if (!revoked) {
return NextResponse.json({ error: "Key not found or already revoked" }, { status: 404 });
}
return NextResponse.json({ success: true, id: params.id, revokedAt: new Date().toISOString() });
}
+118
View File
@@ -0,0 +1,118 @@
import { NextResponse } from "next/server";
import { z } from "zod";
import { isAuthenticated } from "@/shared/utils/apiAuth";
import { issueRegisteredKey, checkQuota, listRegisteredKeys } from "@/lib/db/registeredKeys";
// ─── Validation ───────────────────────────────────────────────────────────────
const issueKeySchema = z.object({
name: z.string().min(1).max(120),
provider: z.string().max(80).optional().default(""),
accountId: z.string().max(120).optional().default(""),
idempotencyKey: z.string().max(256).optional(),
expiresAt: z.string().datetime({ offset: true }).optional(),
dailyBudget: z.number().int().positive().optional(),
hourlyBudget: z.number().int().positive().optional(),
});
// ─── GET /api/v1/registered-keys ─────────────────────────────────────────────
/**
* List registered keys (masked no raw key material returned after creation).
* Optional query params: ?provider=&accountId=
*/
export async function GET(request: Request) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
const { searchParams } = new URL(request.url);
const provider = searchParams.get("provider") ?? undefined;
const accountId = searchParams.get("accountId") ?? undefined;
try {
const keys = listRegisteredKeys({ provider, accountId });
return NextResponse.json({ keys, total: keys.length });
} catch (err) {
console.error("[registered-keys] GET failed:", err);
return NextResponse.json({ error: "Failed to list registered keys" }, { status: 500 });
}
}
// ─── POST /api/v1/registered-keys ────────────────────────────────────────────
/**
* Issue a new registered key.
*
* Checks provider + account quotas before issuing.
* Returns the raw key ONCE it is never stored in plain text.
* Subsequent fetches will only return the masked prefix.
*/
export async function POST(request: Request) {
if (!(await isAuthenticated(request))) {
return NextResponse.json({ error: { message: "Authentication required" } }, { status: 401 });
}
let body: unknown;
try {
body = await request.json();
} catch {
return NextResponse.json({ error: "Invalid JSON body" }, { status: 400 });
}
const parsed = issueKeySchema.safeParse(body);
if (!parsed.success) {
return NextResponse.json({ error: parsed.error.flatten() }, { status: 400 });
}
const { provider, accountId } = parsed.data;
// ── Quota check ──
try {
const quota = checkQuota(provider, accountId);
if (!quota.allowed) {
return NextResponse.json(
{ error: quota.errorMessage, errorCode: quota.errorCode },
{ status: 429 }
);
}
} catch (err) {
console.error("[registered-keys] quota check failed:", err);
return NextResponse.json({ error: "Quota check failed" }, { status: 500 });
}
// ── Issue ──
try {
const result = issueRegisteredKey(parsed.data);
if ("idempotencyConflict" in result) {
return NextResponse.json(
{
error: "Idempotency key already used",
errorCode: "IDEMPOTENCY_CONFLICT",
existing: result.existing,
},
{ status: 409 }
);
}
const { rawKey, ...keyMeta } = result;
return NextResponse.json(
{
key: rawKey, // ← shown ONCE only
keyId: keyMeta.id,
keyPrefix: keyMeta.keyPrefix,
name: keyMeta.name,
provider: keyMeta.provider,
accountId: keyMeta.accountId,
expiresAt: keyMeta.expiresAt,
createdAt: keyMeta.createdAt,
warning: "Store this key securely — it will not be shown again.",
},
{ status: 201 }
);
} catch (err) {
console.error("[registered-keys] issue failed:", err);
return NextResponse.json({ error: "Failed to issue key" }, { status: 500 });
}
}
+103
View File
@@ -0,0 +1,103 @@
/**
* GET /api/v1/search/analytics
*
* Returns search request statistics from call_logs (request_type = 'search').
* Includes provider breakdown, cache hit rate, cost summary, and error count.
*/
import { NextResponse } from "next/server";
import { getDbInstance } from "@/lib/db/core";
import { enforceApiKeyPolicy } from "@/shared/utils/apiKeyPolicy";
export async function GET(req: Request) {
const policy = await enforceApiKeyPolicy(req, "analytics");
if (policy.rejection) return policy.rejection;
try {
const db = getDbInstance();
// Total search requests
const totalRow = db
.prepare(`SELECT COUNT(*) as cnt FROM call_logs WHERE request_type = 'search'`)
.get() as { cnt: number };
const total = totalRow?.cnt ?? 0;
// Today's searches (UTC date)
const todayStart = new Date();
todayStart.setUTCHours(0, 0, 0, 0);
const todayRow = db
.prepare(
`SELECT COUNT(*) as cnt FROM call_logs WHERE request_type = 'search' AND timestamp >= ?`
)
.get(todayStart.toISOString()) as { cnt: number };
const today = todayRow?.cnt ?? 0;
// Errors
const errRow = db
.prepare(
`SELECT COUNT(*) as cnt FROM call_logs WHERE request_type = 'search' AND (status >= 400 OR error IS NOT NULL)`
)
.get() as { cnt: number };
const errors = errRow?.cnt ?? 0;
// Avg duration
const durRow = db
.prepare(
`SELECT AVG(duration) as avg FROM call_logs WHERE request_type = 'search' AND duration > 0`
)
.get() as { avg: number | null };
const avgDurationMs = Math.round(durRow?.avg ?? 0);
// Per-provider breakdown (provider column stores search provider id)
const provRows = db
.prepare(
`SELECT provider, COUNT(*) as cnt
FROM call_logs WHERE request_type = 'search'
GROUP BY provider ORDER BY cnt DESC`
)
.all() as Array<{ provider: string; cnt: number }>;
// Cost per search provider (matching searchRegistry.ts rates)
const COST_PER_QUERY: Record<string, number> = {
"serper-search": 0.001,
"brave-search": 0.003,
"perplexity-search": 0.005,
"exa-search": 0.01,
"tavily-search": 0.004,
};
const byProvider: Record<string, { count: number; costUsd: number }> = {};
let totalCostUsd = 0;
for (const row of provRows) {
const cost = (COST_PER_QUERY[row.provider] ?? 0.001) * row.cnt;
byProvider[row.provider] = { count: row.cnt, costUsd: cost };
totalCostUsd += cost;
}
// Cached: very fast responses (< 5ms) indicate cache hits
const cachedRow = db
.prepare(
`SELECT COUNT(*) as cnt FROM call_logs
WHERE request_type = 'search' AND duration > 0 AND duration < 5`
)
.get() as { cnt: number };
const cached = cachedRow?.cnt ?? 0;
const cacheHitRate = total > 0 ? Math.round((cached / total) * 100) : 0;
return NextResponse.json({
total,
today,
cached,
errors,
totalCostUsd,
byProvider,
cacheHitRate,
avgDurationMs,
last24h: [],
});
} catch (err: unknown) {
const msg = err instanceof Error ? err.message : String(err);
console.error("[/api/v1/search/analytics]", msg);
return NextResponse.json({ error: "Internal server error" }, { status: 500 });
}
}
+5
View File
@@ -69,6 +69,11 @@ export default function LoginPage() {
router.refresh();
} else {
const data = await res.json();
// (#521) If no password is set, redirect to onboarding instead of showing an error
if (data.needsSetup) {
router.push("/dashboard/onboarding");
return;
}
setError(data.error || t("invalidPassword"));
}
} catch (err) {
+11
View File
@@ -281,6 +281,7 @@
"failedUpdatePermissionsRetry": "Failed to update permissions. Please try again.",
"unknownProvider": "unknown",
"copyMaskedKey": "Copy masked key",
"keyOnlyAvailableAtCreation": "Full key available only at creation time — copy it when you first create the key",
"modelsCount": "{count, plural, one {# model} other {# models}}",
"lastUsedOn": "Last: {date}",
"editPermissions": "Edit permissions",
@@ -437,6 +438,11 @@
"antigravityStep2Prefix": "2. Add",
"antigravityStep2Suffix": "to your hosts file as 127.0.0.1.",
"antigravityStep3": "3. Open Antigravity and requests will be proxied.",
"mitmHowWorksDesc": "{toolName} sends requests to its provider endpoint. MITM intercepts and redirects them to OmniRoute.",
"mitmStep1": "1. Start MITM to route requests through OmniRoute.",
"mitmStep2Prefix": "2. Add",
"mitmStep2Suffix": "to your hosts file as 127.0.0.1.",
"mitmStep3": "3. Open {toolName} and requests will be proxied.",
"sudoPasswordRequiredTitle": "Sudo Password Required",
"sudoPasswordHint": "Administrator password is required to modify hosts file and system proxy settings.",
"enterSudoPassword": "Enter sudo password",
@@ -1432,6 +1438,11 @@
"compatDeveloperShort": "Developer role",
"compatDoNotPreserveDeveloper": "Do not preserve developer role",
"compatBadgeNoPreserve": "No preserve",
"compatProtocolLabel": "Client request protocol",
"compatProtocolHint": "These options apply when OmniRoute detects this request shape (OpenAI Chat, Responses API, or Anthropic Messages).",
"compatProtocolOpenAI": "OpenAI Chat Completions",
"compatProtocolOpenAIResponses": "OpenAI Responses API",
"compatProtocolClaude": "Anthropic Messages",
"modelId": "Model ID",
"customModelPlaceholder": "e.g. gpt-4.5-turbo",
"loading": "Loading...",
+5
View File
@@ -1432,6 +1432,11 @@
"compatDeveloperShort": "Developer 角色",
"compatDoNotPreserveDeveloper": "不保留 developer 角色",
"compatBadgeNoPreserve": "不保留",
"compatProtocolLabel": "客户端请求协议",
"compatProtocolHint": "以下选项在 OmniRoute 识别到该请求形态(OpenAI Chat、Responses API 或 Anthropic Messages)时生效。",
"compatProtocolOpenAI": "OpenAI Chat Completions",
"compatProtocolOpenAIResponses": "OpenAI Responses API",
"compatProtocolClaude": "Anthropic Messages",
"modelId": "模型 ID",
"customModelPlaceholder": "例如:gpt-4.5-turbo",
"loading": "正在加载...",
+137
View File
@@ -0,0 +1,137 @@
/**
* Node.js-only instrumentation logic.
*
* Separated from instrumentation.ts so that Turbopack's Edge bundler
* does not trace into Node.js-only modules (fs, path, os, better-sqlite3, etc.)
* and emit spurious "not supported in Edge Runtime" warnings.
*/
function getRandomBytes(byteLength: number): Uint8Array {
const bytes = new Uint8Array(byteLength);
globalThis.crypto.getRandomValues(bytes);
return bytes;
}
function toBase64(bytes: Uint8Array): string {
return btoa(String.fromCharCode(...bytes));
}
function toHex(bytes: Uint8Array): string {
return Array.from(bytes, (byte) => byte.toString(16).padStart(2, "0")).join("");
}
async function ensureSecrets(): Promise<void> {
let getPersistedSecret = (_key: string): string | null => null;
let persistSecret = (_key: string, _value: string): void => {};
try {
({ getPersistedSecret, persistSecret } = await import("@/lib/db/secrets"));
} catch (err: unknown) {
const msg = err instanceof Error ? err.message : String(err);
console.warn(
"[STARTUP] Secret persistence unavailable; falling back to process-local secrets:",
msg
);
}
if (!process.env.JWT_SECRET || process.env.JWT_SECRET.trim() === "") {
const persisted = getPersistedSecret("jwtSecret");
if (persisted) {
process.env.JWT_SECRET = persisted;
console.log("[STARTUP] JWT_SECRET restored from persistent store");
} else {
const generated = toBase64(getRandomBytes(48));
process.env.JWT_SECRET = generated;
persistSecret("jwtSecret", generated);
console.log("[STARTUP] JWT_SECRET auto-generated and persisted (random 64-char secret)");
}
}
if (!process.env.API_KEY_SECRET || process.env.API_KEY_SECRET.trim() === "") {
const persisted = getPersistedSecret("apiKeySecret");
if (persisted) {
process.env.API_KEY_SECRET = persisted;
} else {
const generated = toHex(getRandomBytes(32));
process.env.API_KEY_SECRET = generated;
persistSecret("apiKeySecret", generated);
console.log(
"[STARTUP] API_KEY_SECRET auto-generated and persisted (random 64-char hex secret)"
);
}
}
}
export async function registerNodejs(): Promise<void> {
await ensureSecrets();
const { initConsoleInterceptor } = await import("@/lib/consoleInterceptor");
initConsoleInterceptor();
const [
{ initGracefulShutdown },
{ initApiBridgeServer },
{ startBackgroundRefresh },
{ getSettings },
] = await Promise.all([
import("@/lib/gracefulShutdown"),
import("@/lib/apiBridgeServer"),
import("@/domain/quotaCache"),
import("@/lib/db/settings"),
]);
initGracefulShutdown();
initApiBridgeServer();
startBackgroundRefresh();
console.log("[STARTUP] Quota cache background refresh started");
try {
const [{ setCustomAliases }, { setDefaultFastServiceTierEnabled }] = await Promise.all([
import("@omniroute/open-sse/services/modelDeprecation.ts"),
import("@omniroute/open-sse/executors/codex.ts"),
]);
const settings = await getSettings();
if (settings.modelAliases) {
const aliases =
typeof settings.modelAliases === "string"
? JSON.parse(settings.modelAliases)
: settings.modelAliases;
if (aliases && typeof aliases === "object") {
setCustomAliases(aliases);
console.log(
`[STARTUP] Restored ${Object.keys(aliases).length} custom model alias(es) from settings`
);
}
}
const persisted =
typeof settings.codexServiceTier === "string"
? JSON.parse(settings.codexServiceTier)
: settings.codexServiceTier;
if (typeof persisted?.enabled === "boolean") {
setDefaultFastServiceTierEnabled(persisted.enabled);
console.log(
`[STARTUP] Restored Codex fast service tier: ${persisted.enabled ? "on" : "off"}`
);
}
} catch (err: unknown) {
const msg = err instanceof Error ? err.message : String(err);
console.warn("[STARTUP] Could not restore runtime settings:", msg);
}
try {
const { initAuditLog, cleanupExpiredLogs } = await import("@/lib/compliance/index");
initAuditLog();
console.log("[COMPLIANCE] Audit log table initialized");
const cleanup = cleanupExpiredLogs();
if (cleanup.deletedUsage || cleanup.deletedCallLogs || cleanup.deletedAuditLogs) {
console.log("[COMPLIANCE] Expired log cleanup:", cleanup);
}
} catch (err: unknown) {
const msg = err instanceof Error ? err.message : String(err);
console.warn("[COMPLIANCE] Could not initialize audit log:", msg);
}
}

Some files were not shown because too many files have changed in this diff Show More