Announcements, tutorials, model analysis, and product updates from AIHubMix.
TutorialNewsOpinionAnnouncementChangelog
Migrating from Claude Haiku 4.5 to 5.5: Five 400 Errors and the Quiet ChangesChanging claude-haiku-4-5 to claude-haiku-5-5 is the smallest part of this migration. Five request patterns that worked on Haiku 4.5 now return a 400 error, and several more changes fail no request but alter what you get back, what it costs, or how the model behaves inside an agent. Anthropic says existing Haiku 4.5 prompts should work well on Haiku 5.5 without changes. The request code around those prompts is a different story. This post lists each problem as you'll meet it: what you'll see,
Claude Haiku 5.5 Pricing: The 100K Line Behind the 90% CutClaude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, a tenth of Haiku 4.5's $1 and $5. That holds for prompts up to 100,000 tokens. Above that line, the whole request is billed at $0.50 and $2.50, which is half of Haiku 4.5's price rather than a tenth. Anthropic puts the typical saving at around 75%, not 90%, and the difference comes from three things this post works through: where your prompts fall relative to the 100K line, how many thinking tokens the d
Claude Haiku 5.5 Effort Levels: Medium Is the Default, Low Is Often EnoughClaude Haiku 5.5 is the first Haiku with an effort setting, and the setting moves the bill more than anything else you control. In Artificial Analysis's independent runs, Haiku 5.5 at max effort scored 43 on its Intelligence Index and at low effort scored 29. Max also used 440 million output tokens to get through the index; low used 32 million. So the short answer: leave most work at the default, medium. Drop high-volume, simple routes to low. Raise knowledge work and strict instruction follow
Claude Haiku 5.5 vs Haiku 4.5: What Ten Cents Now BuysClaude Haiku 4.5 scored 0.0% on Terminal-Bench 4.0. Claude Haiku 5.5 scores 39.2%, and it costs a tenth as much per token: $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Haiku 4.5. That is the upgrade in one line. The small Claude model went from "fine for classification, not for agents" to a usable subagent, and its price now matches OpenAI's GPT-6 Luna to the cent. Anthropic released it on October 7, 2026 as Claude Haiku 5.5, model ID claude-haiku-5
Page Assist in Practice: Four Everyday Browser Workflows with Local and Cloud ModelsMost AI chat tools live in their own tab. You copy text out of a page, paste it into a chat window, and copy the answer back. Page Assist removes that round trip: it is an open-source browser extension that opens a model in a sidebar next to whatever page you are reading, plus a full-tab web UI for longer conversations. It started as a front end for local models running in Ollama, and that is still the default. But it also accepts any OpenAI-compatible endpoint, which means the same sidebar can
Migrating to GPT-6.1 Sol: 9 Things That Can Go WrongMoving from GPT-6 Sol to 6.1 Sol looks like a one-line change, and the price is the same. But there are a few breaking changes and some shifts in behavior, so changing only the model name can get you errors, a surprise bill, or an agent that behaves differently. OpenAI's GPT-6 migration guide covers most of the official changes. This post adds the things that tend to bite in practice. They're ordered from "fails loudly" to "fails quietly." 1. reasoning_effort: "none" returns a 400 What you'l
What GPT-6.1 Sol Really Costs: Beyond the $2 / $10 Price TagGPT-6.1 Sol has the same list price as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. Your actual bill depends on four other things: how often you hit the cache, whether you cross 272K input tokens, which service tier you use, and how many reasoning tokens the model burns. This post goes through each one. The full price list Standard rates (per 1M tokens) GPT-6.1 SolGPT-6 SolGPT-6 AstraGPT-6 LunaInput$2.00$2.00$10.00$0.10Cached input$0.10$0.20$1.00—Cache writes$
Choosing a Reasoning Effort for GPT-6.1 Sol: low to maxGPT-6.1 Sol has five reasoning effort levels: low, medium (the default), high, xhigh, and max. The big change from GPT-6 Sol is that none and minimal are gone, so low is now the floor. This one setting drives both latency and cost. You never see reasoning tokens, but you pay for them at the output rate ($10 per million for 6.1 Sol on AIHubMix), and they take up room in the context window. Pick the wrong level and you can easily pay several times more than you need to. What each level is for
GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches AstraOpenAI shipped GPT-6.1 Sol at DevDay on September 29, 2026. The timing is striking: GPT-6 Sol had been out for exactly one week (it launched September 22 alongside Luna), and the flagship GPT-6 Astra was less than a month old. The official model page sums up the pitch in one line: "Near-Astra performance for complex work at a lower cost." Concretely, that means agentic coding, computer use, and professional tasks at roughly Astra quality, for one fifth of Astra's per-token price. This post ans
AIHubMix Adds Claude Sonnet 5.5: An Upgrade for Everyday Work at Unchanged Token PricesA 1-million-token context window, input at $2 and output at $10 per million tokens, with browser access and a shared gateway to multiple model providers. AIHubMix now offers Claude Sonnet 5.5, Anthropic's latest Sonnet model. Users can access it through AIHubMix Playground or compatible applications for writing, organizing information, preparing presentation content, and other everyday work. Compared with Sonnet 5, Anthropic reports faster output generation, clearer writing, and more efficient
Xiaomi MiMo v2.6 Pro Quick Start: Three Steps from Browser to CodeThe quick answer: open https://playground.aihubmix.com/ in your browser, type a question, press send. No installing, no coding, and free trial calls without entering a card. Key points * The model is mimo-v2.6-pro, made by Xiaomi, released 22 September 2026 * You can use it in a browser for free before deciding anything * It reads text, images, video and audio * It costs $0.48 per million input tokens, $0.96 per million output — a fraction of a cent for a typical question * The older mim
How to Use Xiaomi MiMo v2.6 Pro on AIHubMix: Setup GuideIn short: point the official openai Python SDK at https://aihubmix.com/v1, use model ID mimo-v2.6-pro, and authenticate with your AIHubMix key. The API is OpenAI-compatible (callable in the same format as OpenAI's), so no special client is required. Setup takes about five minutes. Key points * OpenAI-compatible API — change base_url, nothing else * Paid mimo-v2.6-pro · free xiaomi-mimo-v2.6-pro-free [1] [2] * $0.48/M input · $0.96/M output [1] * Free tier 5 requests/min · 100/day · 1M tok
How to Call Jev on AiHubMix: A Structured Classification TutorialShort answer: jev-1.13 does not generate text. You POST a piece of text plus a set of named questions to https://aihubmix.com/v1/systemone, and you get back typed answers keyed by your question names — a category, a score, or a probability. Nothing to parse. Below is a working request, the full response format, and the one design mistake that will quietly cost you accuracy. What Jev is for Use it when you need a judgment, not a paragraph: routing a support ticket, scoring severity, flagging u
DeepSeek V4.1 Flash API Pricing Compared: AIHubMix 30% Off vs OpenRouterDeepSeek V4.1 Flash is now available on AIHubMix and OpenRouter. It supports text and image inputs, a 1M-token context window, tool calling, structured outputs, and agent-oriented workflows. AIHubMix currently offers a limited-time 30% discount on selected provider routes, valid through September 27, 2026. OpenRouter lists the price at $0.13 per million input tokens and $0.52 per million output tokens, with cache reads at $0.0026 per million tokens. DeepSeek V4.1 Flash API pricing at a glance
Why did Gemini turn Chinese for no reason?Gemini may occasionally produce Chinese in its reasoning or final response even when you expect English. This usually points to temporary multilingual language drift. It does not, by itself, mean that Gemini has switched to a separate Chinese model. The short version: inspect the full conversation for Chinese language signals, restate the desired output language, and do not use the stream setting as a language fix. What Gemini language drift means Language drift occurs when a model begins in
Jev Explained: How to Add Fast, Typed Decisions to an AI AgentJev is best understood as a decision layer for software. It reads text or structured state and returns predefined classifications, scores, and yes-or-no probabilities. It does not write an answer for the user. That narrower interface makes it relevant to high-volume routing, triage, verification, and guardrail steps inside AI agents. The practical pattern is simple: let Jev make frequent, reversible judgments; let business code enforce policy; escalate uncertain or consequential cases to a capa
How to Create an AI Product Ad with Claude Code and AIHubMixWith one clear product image, Claude Code, and the AIHubMix API, you can generate character references in Seedream, register them as virtual-person assets, and use Seedance to create an AI product ad. This tutorial is based on a real aloe-mist campaign and includes the project structure, configuration, commands, and complete video prompt. The target output is a 20-second, 720p, 9:16 vertical ad with an AI-generated person and setting. Claude Code and AIHubMix power the full workflow. Claude Co
Amp Reopens BYOK After a Year: Connecting and Configuring AIHubMixAmp has reopened BYOK (Bring Your Own Key), and AIHubMix now supports it. This guide uses Anthropic's Claude models as the example and walks through pointing Amp at the AIHubMix API via Custom URL and Model Routing — covering the API key, model name mapping, and project setup. If you're searching for "Amp BYOK setup", "connect Amp to AIHubMix", "AIHubMix API key", "Amp Custom URL", or "Amp Model Routing", the steps below cover it. What is Amp good for? * Multi-model workflows: assign differ
How to Use a Real Face in Seedance 2.5 with AIHubMix Real-Human AssetsAI video becomes much more useful when the person on screen stays recognizable from one scene to the next. AIHubMix real-human assets provide a consent-based way to use an approved face in Seedance video generation while keeping authorization, asset processing, and generation as separate, verifiable steps. This guide explains what real-human video is, how the AIHubMix workflow works, why it is useful, and how to create a Seedance 2.5 video from a real face. It also explains the difference betwe
GLM-5.3-Flash Pricing Compared: OpenRouter, Z.ai, and AIHubMixGLM-5.3-Flash pricing on OpenRouter, Z.ai, and AIHubMix, including OpenRouter's 5.5% platform fee and cache-ratio cost examples.
AI Agent Architecture: Model Routing and Tool DiscoveryLearn why production AI agents need two integrations: AIHubMix for real-time model routing and Monid for runtime tool discovery and API access.
Building the Agent Economy: AIHubMix and FluxA Announce Strategic CollaborationAt AIHubMix, we believe the next phase of artificial intelligence will be defined not only by what models can understand, but also by what AI agents can responsibly accomplish. That is why we are announcing a strategic collaboration with FluxA, an agent-native payment infrastructure provider focused on wallets, identity, authorization, service monetization, and payments for autonomous systems. Together, AIHubMix and FluxA are working to connect two essential layers of the Agent Economy: access
Tell Your Agent One Sentence, Get 800+ Models"Agent-readable access on the AIHubMix AI gateway: agents.md, llms.txt for 850+ models, an Agent Skill, MCP, and deep links for Claude Code, Codex, Cursor.
GPT-5.6 Sol API Pricing: 50% Off on AIHubMix and OpenRouterGPT-5.6 Sol is currently 50% off on AIHubMix and OpenRouter, but OpenRouter adds a 5.5% Pay-as-you-go platform fee. Compare the real final cost.
GLM-5.3 Pricing Compared: OpenRouter, Z.ai, and AIHubMixGLM-5.3 has quickly become one of the most interesting models for coding and long-horizon agent work. But the price you pay depends heavily on where you access it. The official Z.ai API and OpenRouter currently list the model at $1.40 per million input tokens and $4.40 per million output tokens. AIHubMix offers a separate coding-glm-5.3 preview route at $0.06 per million input tokens and $0.22 per million output tokens. That is a dramatic difference. It is also not a simple apples-to-apples co
GLM-5.3 Hands-on Guide: Always-on Thinking, Three Effort Levels, and the API Support MatrixAn August 2026 guide to calling GLM-5.3: always-on thinking with three reasoning_effort levels, reasoning summaries, parallel tool calls, structured output, and automatic caching — with tested examples for the AIHubMix Chat, Responses, and Messages APIs.
DeepSeek V4 Pro (0813): Thinking Passback & 3-API MatrixDeepSeek V4 Pro (0813) hands-on guide: thinking toggle and reasoning_effort levels, mandatory thinking-history passback, tools, caching, and a 3-API matrix.
DeepSeek V4 Flash Was Degraded Today. Here’s Why Multi-Provider Failover MattersOn August 4, 2026, DeepSeek’s official status page recorded two API degraded-performance incidents. The first incident lasted 1 hour and 18 minutes, from 02:02 to 03:20 UTC, and affected DeepSeek V4 Flash, V4 Pro, and Expert Mode. The second incident lasted 36 minutes, from 03:43 to 04:20 UTC, and affected the DeepSeek V4 Flash API. Both incidents have since been resolved. OpenCode also reported that DeepSeek Flash was experiencing capacity issues due to unprecedented demand. However, DeepSeek
June 2026 Release Spotlight: ~20 New ModelsIn June 2026 AIHubMix added ~20 models (glm-5.2, minimax-m3, qwen3.7-plus, kimi-k2.7-code, Kling video) plus LLM Router, Mapping & Fallback, CLI, backup domain.
Global Acceleration: 75% Lower Latency, 99.99% AvailabilityAIHubMix runs a self-built acceleration network: 75% lower latency, 60% less fluctuation, 99.99% availability, minute-level probes, automatic failover.
OpenAI Compatible Interface Upgrade: Deep Support for ClaudeAIHubMix upgraded its OpenAI-compatible API for Claude: interleaved thinking with no extra parameters, prompt caching, and Anthropic beta feature support.
GPT-5.6 Is Live: Prompt Caching Billing Changes ExplainedJuly 2026: GPT-5.6 gpt-5.6-sol / terra / luna on AIHubMix: 1.05M context, 1.25x cache writes, prompt_cache_key, explicit breakpoints, vs Claude caching.
July 2026 Release Spotlight: ~30 New Models and Media APIsAIHubMix added ~30 models in July 2026, including claude-opus-5, GPT-5.6, kimi-k3 and qwen3.8-max-preview, plus media generation, 3D generation and MCP Server.
Free AI Models on AIHubMixFree AI Models: The Ultimate Guide to Building with Zero-Cost AI in 2026 on AIHubMix
Kimi K3 Hands-On Guide: New Parameters & API Support MatrixJuly 2026 Kimi K3 guide: reasoning_effort max, thinking history, dynamic tool loading, structured output, auto caching, partial prefix, and vision inputs.
Claude Opus 4.7 New Parameters GuideThis article covers two key changes to reasoning control in Claude Opus 4.7, along with complete usage instructions for both the AIHubmix native API and the Chat unified interface. See also: Anthropic official announcement and model change log.