Capability change announced Apr 27, 2026
OpenRouter documents a capability change: “Insights — OpenRouter Blog ## Agent Frameworks Compared: Tool-Calling Schema Handling OpenAI, Anthropic, and Google each use a different request and…”
- Hosting route
- First-party API
- Affected scope
- Not stated in the notice
- Announced
- Apr 27, 2026
- First seen by ModelClock
- Oct 3, 2026
What the provider published
openrouter.ai ↗Insights — OpenRouter Blog ## Agent Frameworks Compared: Tool-Calling Schema Handling OpenAI, Anthropic, and Google each use a different request and response shape for the same tool. Agent frameworks handle that difference in different places. Some translate one definition into each provider's format, some are native to a single provider, and some hand the question to a connector underneath. This article compares six frameworks on schema definition, translation, and MCP support, then shows how OpenRouter normalizes tool calling at the API layer so a model swap is a change to one string. Entry date: 2026-10-02 ## LangChain vs CrewAI: Orchestration Compared to OpenRouter-Native Routing Multi-model orchestration is three layers. Workflow orchestration is planning, state, memory, and delegation, and LangGraph and CrewAI are built for it. Model routing is choosing a model per call and falling back when it fails, and provider routing is choosing which provider serves that model. OpenRouter does the second two. This article separates the layers, shows the same two-step pipeline in direct OpenRouter calls and in LangChain, and describes when to add a framework, when to use our Agent SDK, and how to put OpenRouter underneath LangGraph or CrewAI. Entry date: 2026-10-02 ## Confidence Thresholds for Model Escalation Routing Confidence-based escalation keeps most requests on a cheap model and sends only the ones it is unsure about to a stronger one. This guide covers forcing a numeric confidence field with structured outputs, setting the threshold from error rates on your own traffic, tuning it against accuracy, cost, and latency, routing the escalation in your code, and re-tuning it after launch. Entry date: 2026-10-01 ## Cost vs. Quality Tradeoff Framework for Agent Models The cheapest model that clears your quality bar is usually not the model at the top of a leaderboard. This framework sets the bar for one task, measures cost per quality point across a cheap, a mid-tier, and a frontier model on your own examples, and picks the cheapest one that clears the bar with margin. It uses live catalog prices and the cost each OpenRouter response reports. Entry date: 2026-10-01 ## Image-to-Video AI Models Compared: Cost, Resolution, and Control If you already have the image a video should start from, the model choice comes down to what has to happen after that frame. This post compares the Veo 3.1, Seedance, Kling, and Grok Imagine Video lines on duration, resolution, first-frame and last-frame control, generated audio, reference inputs, and current per-second price, then shows how to submit an image-to-video job with the TypeScript SDK. Entry date: 2026-09-29 ## Best Embedding Models in 2026 An embedding model decides what your retrieval system can find. We shortlisted the embedding models in our catalog for English RAG, multilingual retrieval, code search, text-and-image retrieval, and low-cost indexing, sent live requests to each one, and recorded their prices, context windows, and default vector sizes so you can pick a candidate and test it on your own data. Entry date: 2026-09-23 ## Is Jev as Accurate as Frontier Models at Classification? Claude Opus 5 leads OpenRouter's classification task ranking by spend. We sent the same 3,080 Banking77 utterances to it and to Jev 1.13 through the Decisions API. Opus scored 84.4% to Jev's 81.0%, and Jev answered in 175 ms at $0.11 per thousand requests against 2.3 seconds and $2.42 for Opus. Entry date: 2026-09-22 ## What Is Nemotron 3.5 Lightning Nemotron 3.5 Lightning is NVIDIA's open-weight 30B mixture-of-experts model with about 3B active parameters per token, built for the high-volume execution calls in an agent run. This post covers what the architecture means, how the model compares with Nemotron 3 Ultra, what context and features each OpenRouter endpoint exposes, and how to call it with structured outputs. Entry date: 2026-09-22 ## Two Hours of Work That Takes a Week For Descript, testing a new model was a couple of hours of work and a week of waiting. The team removed the waiting. Evaluations run in an hour or two now, and nobody has to ask an engineer for time. This post covers the queue, what replaced it, and what changed once evaluation ran several times a week. Entry date: 2026-09-21 ## What Is Jev? TypeSafe's Decision Model Explained for Developers Jev reads text and returns typed answers with probabilities instead of prose. Here's what that means, three live API calls, how to read the numbers, and how to call it with an OpenRouter key. Entry date: 2026-09-21 ## Image Generation Models Compared: Cost, Edit, Quality Image models are priced per token, per megapixel, or per image, so their listed rates do not compare. We sent the same prompt to 20 of them through the Image API, read usage.cost off each response, and tested text rendering, reference-image editing, seeds, and text-plus-image replies. Entry date: 2026-09-18 ## Does DeepSeek V4 Have Vision? DeepSeek V4 is a family of models, and only some of them accept images. This guide maps every V4 slug on our catalog to its input modalities, shows how to send an image to the two that read images, and shows how to put a vision model in front of the text-only ones. Entry date: 2026-09-16 ## Zero Data Retention (ZDR): What It Means for AI APIs ZDR means an AI provider processes your prompt, returns a response, and doesn't store either one afterward. It's a retention guarantee, not a universal privacy policy. This page explains the boundary, compares ZDR with related controls, and shows how to enforce it at the account, guardrail, or request level. Entry date: 2026-09-11 ## OpenRouter Fusion: How It Works and When to Use It Fusion is our compound model. It turns one prompt into a short debate among several models: a panel answers in parallel, a judge maps agreement and disagreement, and the calling model writes the final answer. You trade some speed and tokens for quality. This is what it does, what it costs, and when to use it. Entry date: 2026-09-10 ## Seedance 2.5 Review: What It's Best At and When to Use It Seedance 2.5 trades resolution for length. It runs to 30 seconds and stops at 720p. We work through what it's best at, what a clip costs at each resolution, how it compares to Seedance 2.0, Wan 3.0, and Veo 3.1 on our own catalog data, and the cases where we'd recommend a different model. Entry date: 2026-09-09 ## GPT 5.6 Discounts & Jevons Paradox OpenAI introduced large discounts on their new Terra and Luna models from July 27th through August 14th. What impact did these discounts have on token volumes, total spend, and the competition? Entry date: 2026-08-25 ## Governing AI Spend Across a Team on OpenRouter Your team's AI bill is the total of every key each engineer holds. OpenRouter gives you six ways to control that spend, from per-key limits to Enterprise workspace budgets. This guide explains what each control caps, what plan it needs, and which ones fit your team. Entry date: 2026-08-06 ## How to Evaluate LLM Provider Performance Across Latency, Throughput, and Uptime The same model behaves differently across provider endpoints. Infrastructure, quantization, load handling, and routing defaults all change the result. Here's how to measure latency, throughput, uptime, and precision, then turn the measurements into a routing policy. Entry date: 2026-07-28 ## Every Modality Through One API Chat, image generation, embeddings, and transcription usually mean four SDKs, four bills, and four auth schemes. On OpenRouter every modality runs through one base URL: you change the model string and the content type, and the same routing controls carry across. Entry date: 2026-07-16 ## Why Use OpenRouter for DeepSeek DeepSeek is one model served by 16 providers, at prices that vary by about 4x and throughput from 4 to 57 tokens per second. Here's what routing that spread through one slug actually buys you, and when going direct is the better call. Entry date: 2026-07-13 ## Choosing the Optimal Image Input Detail Level in LLMs We ran 1,730 visual reasoning questions across 5 models. Dropping image detail to "low" costs real accuracy, and on gpt-5.5 the bill went up too. The lever that reliably cuts cost is reasoning effort. Entry date: 2026-07-07 ## DeepSeek V4 Is Earning Agentic Token Share DeepSeek doubled its token share on OpenRouter in six months. V4 Flash is the model that made it happen, and agentic workloads are driving the surge. Entry date: 2026-06-30 ## The Open Weight Models that Matter: June 2026 A slew of compelling open-weight models have shipped from new players in both China and the US. As of June 2026, these are the four open-weight models that matter the most — and when to reach for each. Entry date: 2026-06-27 ## AI Governance Checklist: Your LLM Architecture Comes First Policy language can't show who called which model or where the audit trail lives. Map your governance checklist to the three routing postures your stack can actually prove. Entry date: 2026-06-22 ## How to Enforce AI Data Residency Without Building Local Infrastructure If your procurement team flagged country of origin, you don't need to build local infrastructure. For API teams, data residency is a routing constraint you enforce in a single request. Entry date: 2026-06-22 ## OpenRouter vs Portkey: Which LLM Gateway for Your Team? OpenRouter routes across providers on credits you buy; Portkey governs the provider keys you already have. Here's how they compare on models, observability, compliance, and price. Entry date: 2026-06-19 ## OpenRouter vs LiteLLM: Which LLM Gateway Fits Your Stack? OpenRouter is a managed gateway; LiteLLM is a self-hosted proxy. Here's how they compare on cost, data residency, routing, and latency. Entry date: 2026-06-19 ## Agentic AI Governance: Your API Key Is a Guardrail Most agentic AI governance frameworks define the rules but never enforce them. Learn how the API routing layer stops runaway spend and model escalation. Entry date: 2026-06-15 ## How OpenRouter Model Routing Works OpenRouter routes every request across 70+ providers, and you control how: provider order, price ceilings, and the fallback chain. Here's how each routing layer works. Entry date: 2026-06-12 ## OpenRouter Reliability & Automatic Failover: How Requests Keep Succeeding Provider failover is on by default. Model fallbacks are opt-in. The two layers recover from different failures; here's how each one works and where failover stops. Entry date: 2026-06-12 ## Dinner is Served Standardizing on one LLM is like everyone at the table ordering their own entree. The case for going family style with your AI, and the OpenRouter data showing teams already do. Entry date: 2026-06-11 ## What Is an LLM Gateway? The Missing Layer Between Your App and AI Models Without an LLM gateway, provider outages become user-facing errors and AI spend stays opaque. Compare the best options by routing, compliance, and setup time. Entry date: 2026-06-11 ## A Robot is Sprinting Towards You: Do You Want it Running on Claude or Grok? A 30-game battle royale across eleven LLMs, $482 of inference, and one finding that should change how you read model benchmarks. Entry date: 2026-06-04 ## GPT-5.5 Price Increase: What It Actually Costs OpenAI doubled per-token prices with GPT-5.5 but the model is less verbose. We measured real usage to see the net cost impact. Entry date: 2026-05-04 ## Opus 4.7's New Tokenizer: What It Actually Costs Anthropic changed the tokenizer in Opus 4.7. We looked at usage that shifted from 4.6 to 4.7 to measure exactly how it affects costs. Entry date: 2026-04-27 ## The 2025 State of AI Report Introducing the 2025 State of AI report, in partnership with a16z. The largest empirical look yet at how developers and organizations use language models in the real world. Entry date: 2025-12-04 ## Is Implicit Caching Prompt Retention? Should customers consider providers that have implicit caching as “ZDR”? Entry date: 2025-10-23
Read from the provider's text; the quote is the provider's exact lines.