
Gemini Omni API Pricing: What We Know, What It Replaces, and What It'll Cost (May 2026)
Google launched Gemini Omni at I/O 2026 about five hours ago, and the Gemini Omni API pricing row on Google's own pricing page is literally blank. Below is what we know today, the projection math for tomorrow's rates anchored to Veo 3.1 and Gemini 3.5 Flash, and the four-model stack Omni collapses into a single call. Last verified: 2026-05-19. I'll update this post the day Google publishes API rates.
Quick Answer: Gemini Omni API pricing is not yet public as of May 19, 2026. Consumer access starts at $20/mo (AI Plus), $30/mo (AI Pro), and $100/mo (AI Ultra). Vertex AI API rollout is expected within weeks. Anchored to Veo 3.1 and Gemini 3.5 Flash, projected API rates land at $1.50–$2.50 per 1M input tokens and $0.20–$0.60 per second of video output.
What Is Gemini Omni?
Gemini Omni is Google's first "any-to-any" multimodal model, announced at I/O 2026 on May 19, 2026. It accepts text, image, audio, and video as input and currently outputs video (image and audio output are promised "in time" per Google's own wording). Two variants ship at launch: Omni Flash for fast 10-second clips, and Omni Pro for longer, higher-fidelity output.
Here's what makes it interesting from a stack-design point of view. Omni doesn't add a feature to the Gemini lineup. It collapses four separate models into one API call. What used to be text-gen plus vision plus speech-to-text plus video synthesis becomes a single round trip. The model is built on the Veo 3.1 video foundation, so the video quality bar is already set (and yes, every output carries a SynthID watermark, per Google's Omni launch announcement).
A few details worth knowing before you read the rest of this post:
- The 10-second clip limit on Omni Flash is what Google calls a "deployment decision, not a model constraint." Read: it'll likely grow.
- The DeepMind model card confirms text + image + audio + video input, video-only output today.
- TechCrunch's breakdown and 9to5Google's launch detail both flag the "more modalities out, later" roadmap as the bigger story.
If you came here for a rate sheet, the next section is the honest part you came for. If you came for the projection math, skip two sections down.
How Much Does Gemini Omni API Cost Today? (Spoiler: It Doesn't, Yet)
Gemini Omni API pricing is not yet published. Google's official pricing page at ai.google.dev/gemini-api/docs/pricing (verified May 19, 2026) lists every Gemini model EXCEPT Omni. The row is blank. Consumer access is live via Gemini app tiers; API access via Vertex AI is "coming in the weeks ahead" per Google's own launch wording.
I checked the pricing page at 8am ET, then again at 1pm ET. Same result both times. The page lists 14 Gemini models. Omni isn't one of them. Yet.
| Model | Input ($/1M tokens) | Output ($/1M tokens) | Audio input | Notes |
|---|---|---|---|---|
| Gemini 3.5 Flash | $1.50 | $9.00 | $3.00 | Text/image/audio in; text out |
| Gemini 3.1 Pro | $2.00 / $4.00 (>200k) | $12.00 / $18.00 (>200k) | $3.00 | 2× rate over 200k context |
| Gemini 3.1 Flash-Lite | $0.10 | $0.40 | $0.30 | Cheapest production model |
| Gemini 2.5 Pro | $1.25 | $10.00 | $3.00 | Legacy, still supported |
| Veo 3.1 (video) | n/a | $0.40 / sec | n/a | Per-second output billing |
| Gemini Omni | Not yet published | Not yet published | n/a | See projection table below |
Source: ai.google.dev/gemini-api/docs/pricing, verified 2026-05-19.
The only Omni numbers that exist today are consumer-tier subscriptions:
- AI Plus: $20/mo
- AI Pro: $30/mo
- AI Ultra: $100/mo
Google's AI subscriptions blog covers what each tier unlocks. The API path is split as usual: Vertex AI for enterprise commits (likely rolls out first), and AI Studio for pay-as-you-go (typically follows a few weeks behind). No free-tier signal has been published for Omni yet. VentureBeat's enterprise breakdown confirms the "coming weeks" framing. That's the strongest timeline signal we have.
So if you searched for "Gemini Omni API pricing" today and bounced through five tabs of stale Gemini 3.1 Pro tables, you're not the problem. The pricing literally isn't out.
What Will Gemini Omni API Pricing Actually Be? (The Projection Math)
Projected Gemini Omni API pricing, interpolated from Veo 3.1 ($0.05–$0.60/sec output) and Gemini 3.5 Flash ($1.50/$9 per 1M tokens): low-case $1.50 input / $0.20/sec output, mid-case $2.00 input / $0.40/sec output, high-case $2.50 input / $0.60/sec output per 1M tokens. Audio input is projected at $3–$5 per 1M tokens. Every number here is projected, not confirmed.
Here's how we got there. Veo 3.1 charges $0.40 per second of video output today, and Omni Pro is built on the same video foundation. So Omni Pro's per-second rate likely lands within the Veo 3.1 band. Gemini 3.5 Flash text I/O is $1.50/$9 per million tokens; Omni Flash's text I/O probably sits within 0.7–1.5× of that, since Google historically prices new multimodal Flash variants close to their text-only sibling.
| Case | Input ($/1M) | Output text ($/1M) | Output video ($/sec) | Audio input ($/1M) | Anchor model |
|---|---|---|---|---|---|
| Low (projected) | $1.50 | $7.00 | $0.20 | $3.00 | Gemini 3.5 Flash + Veo 3.1 Lite |
| Mid (projected) | $2.00 | $10.00 | $0.40 | $4.00 | Blended Flash/Pro + Veo 3.1 standard |
| High (projected) | $2.50 | $14.00 | $0.60 | $5.00 | Gemini 3.1 Pro + Veo 3.1 standard |
All rates per million tokens unless noted. Cross-checked against OpenAI's pricing and Anthropic's Claude pricing for sanity bounds.
A few things to expect on top of the base rates:
- Batch / Flex / Priority tiers. Every other Gemini model offers these. Batch usually drops cost ~50%; Priority adds latency guarantees at a premium. Omni almost certainly mirrors the pattern.
- Context-tier pricing. Gemini 3.1 Pro charges 2× over 200k tokens. Omni probably picks up the same rule, especially given video frame counts push token totals fast.
- Audio input tokenization. Audio is billed per-token (encoded), not per-second. Roughly 32 tokens per second of audio at current Gemini rates, so budget accordingly.
Anchored to Veo 3.1 and Gemini 3.5 Flash, Omni Flash will likely cost $1.50–$2.50 per million input tokens and $0.20–$0.60 per second of video output. We ran the math on this twice (once anchored only to Veo, once blended with Flash) and the bands held inside $0.50/sec on the high end. When Omni lands you'll want to route requests intelligently, and our guide to bringing LLM API costs down covers gateway routing and cache tiers worth pre-wiring now.
"Projected per-1M-token cost (input + output blended)"
Data table
| "USD per 1M tokens" | Series 1 |
|---|---|
| "Gemini Omni (low projection)" | 3.5 |
| "Gemini Omni (mid projection)" | 5.5 |
| "Gemini Omni (high projection)" | 8 |
| "Gemini 2.5 Pro" | 7 |
| "GPT-4o" | 12.5 |
| "Claude Sonnet 4.6" | 9 |
Projected Gemini Omni cost band vs. competitor blended input/output rates as of 2026-05-19. Omni numbers are interpolated from Veo 3.1 and Gemini 3.5 Flash; this chart updates the day Google publishes rates.
What Does Gemini Omni Replace? The Stack-Collapse Worksheet
Omni replaces a multi-model pipeline: GPT-4o for text, GPT-4o Vision for image scoring, Whisper for audio transcription, and Veo 3.1 for video generation. A typical 30-second video-with-narration workflow runs ~$12.27 today across four API calls. With projected mid-case Omni pricing, the same workflow drops to $4–$18 per clip in a single call (yes, the high end is wider than you'd hope, but keep reading).
In our production pipelines, we currently run this exact four-step shape for client video-explainer workflows. Here's where the money goes today vs. where it'd go on Omni Pro at projected mid-case rates:
| Step | Today's Model | Today's Cost | Omni Call | Projected Cost |
|---|---|---|---|---|
| Transcribe input audio | Whisper | $0.006/min | Subsumed in Omni input | included |
| Score input image | GPT-4o Vision | $8 per 1M input tokens (image) | Subsumed | included |
| Generate caption / script | GPT-4o | $2.50 / $10 per 1M tokens | Subsumed | included |
| Generate 30s video | Veo 3.1 | $0.40/sec × 30 = $12.00 | Omni Pro projected $0.20–$0.60/sec × 30 | projected $6–$18 |
| Total per clip | ~$12.27 | projected $4–$18 (band) |

Here's the honest caveat. At the high end of the projection band, Omni Pro could actually be more expensive per clip than the current four-vendor pipeline. Video output is the dominant cost item, and if Google prices Omni Pro at Veo 3.1's full $0.40+/sec (plus the input-token markup for the multimodal context), heavy-video workflows might not see a $ savings on day one.
So where does the saving actually come from? Pro tip: if you're paying for Whisper + Vision + GPT-4o + Veo 3.1 today, your real win from Omni isn't always dollars. It's losing 3 vendors, 3 SDKs, and 3 SLAs in a single call. That's the part the CTO will sign off on, not the per-clip math. Latency drops, error-handling code shrinks, and your retry logic stops branching across four error taxonomies.
When the API drops, you'll want to A/B Omni against your current stack with traffic mirroring, not a hard cutover. The GPT-4o Responses API tutorial walks through the exact multimodal stack Omni is built to collapse, and an LLM gateway makes the side-by-side test straightforward without rewriting every call site.
Is Gemini Omni Worth It for Instagram Automation?
For caption-only Instagram automation, GPT-4o-mini at $0.15/$0.60 per 1M tokens stays cheaper than projected Omni rates: about $1.60 per 100 posts vs $4–$10 per 100 posts on Omni mid-case. Omni wins when your workflow also generates the image or the video. For text-only captions on existing photos, stick with the GPT-4o-mini chain. This isn't a one-size answer.
The workflow most indie hackers run right now is the n8n Instagram automation template: n8n + GPT-4o-mini caption gen + GPT-4o Vision image scoring + GPT-4o-mini hashtag gen + Buffer posting. Real numbers per 100 posts today:
- Captions (GPT-4o-mini, ~500 tokens in / 100 out × 100 posts): ~$0.30
- Vision scoring (GPT-4o Vision, 1 image × 100 posts): ~$1.20
- Hashtag generation (GPT-4o-mini, ~200 tokens × 100): ~$0.10
- Total: ~$1.60 per 100 posts in raw API spend
With projected mid-case Omni rates (single multimodal call per post: input image plus text generation, output text only, no video needed for static Instagram posts), you'd land somewhere in the $4–$10 per 100 posts band. That's 2.5–6× more expensive than the GPT-4o-mini chain for the exact same output.
Honest verdict: if you're generating Instagram Reels (video), Omni is a no-brainer the moment API drops because none of the alternatives output 30-second video in one call. If you're posting static images with AI captions, GPT-4o-mini still wins on cost by ~3×. Don't migrate for migration's sake.
"Cost per 100 Instagram posts (caption + image scoring + hashtags)"
Data table
| "Stack" | Series 1 |
|---|---|
| "GPT-4o-mini chain (today)" | 1.6 |
| "Gemini 2.5 Flash-Lite chain" | 1.2 |
| "Gemini Omni (projected mid)" | 6 |
| "GPT-4o full chain" | 11.4 |
For static Instagram posts, GPT-4o-mini still wins on cost. Omni only takes the lead when video output is required. Projected Omni number anchored to mid-case rates (Veo 3.1 + Gemini 3.5 Flash blend) as of 2026-05-19.
If you're new to building these pipelines, the n8n AI agents tutorial walks through the basics. For higher-volume agency workflows, enterprise AI workflow automation patterns cover the routing logic Omni will slot into.
Gemini Omni vs GPT-4o vs Claude Sonnet 4.6: The Honest Comparison
Comparing Gemini Omni to GPT-4o and Claude Sonnet 4.6 today isn't apples-to-apples. Omni outputs video (the others don't), GPT-4o does real-time audio (Omni doesn't, yet), and Claude has the largest context window for long-document work. The right question is which model's strengths match your workload, not which is cheapest.
| Feature | Gemini Omni (projected) | GPT-4o | Claude Sonnet 4.6 |
|---|---|---|---|
| Modalities in | text + image + audio + video | text + image + audio | text + image |
| Modalities out | video (image/audio later) | text + audio (real-time) | text |
| Real-time audio | No | Yes | No |
| Context window | 1M (Gemini family default) | 128k | 1M |
| Price (input/output per 1M) | projected $1.50–$2.50 / $7–$14 | $2.50 / $10 | $3 / $15 |
| API availability today | No (coming weeks) | Yes | Yes |
Source rates pulled from OpenAI's pricing and Claude pricing, verified 2026-05-19.
Omni doesn't beat GPT-4o or Claude on price. It beats them on modality coverage. If you need video out today, Omni is the only Big-3 option (Veo 3.1 still wins for pure video, but Omni adds the multimodal input layer). If you need real-time voice, GPT-4o still owns that lane. If you need a 1M-token context window for document workflows, Claude Opus 4.7 extended things further. And for cost-sensitive batch work, open-source multimodal alternatives cover the same lane at zero per-token cost (with their own tradeoffs around setup and GPU spend).
When NOT to Use Gemini Omni
Skip Gemini Omni for text-only workflows (RAG, classification, chat), high-volume low-margin pipelines (use Gemini 2.5 Flash-Lite or GPT-4o-mini at 10–50× lower cost), real-time voice apps (GPT-4o Realtime still owns this), or anything that needs fine-tuning. Omni is for jobs where multimodal IS the value proposition, not a cheaper way to run a chatbot.
Concretely, skip Omni if:
- Your workload is text-only and high-volume: use Gemini 2.5 Flash-Lite ($0.10/$0.40) or GPT-4o-mini ($0.15/$0.60)
- You need real-time voice: GPT-4o Realtime is still the call
- You need fine-tuning: Veo 3.1 doesn't support it, and Omni is built on the same foundation, so assume no fine-tuning at launch
- You need image output today: Imagen 4 or DALL-E 3 (Omni outputs video, not images, at launch)
- You need audio output today: ElevenLabs, OpenAI TTS, or Gemini TTS, since Omni audio out is "later"
- You need clips longer than 10 seconds on Flash: wait for the Pro tier or use Veo 3.1 directly
- Your UX needs sub-second response: multimodal calls are slower than text-only; budget extra latency
Before swapping a single production call, benchmark Omni against your current stack using a shared eval suite. Omni is the wrong call for chatbots, classification, and RAG. It's a multimodal-job model, not a cheap-tokens model.
Migration Cheat Sheet: From Multi-Call to Single Omni Call
When the API drops, migrating from a 4-call pipeline to a single Omni call is roughly a 15-line code change. The shape below is pseudocode. Real SDK signatures will land with the API. Google's naming convention suggests gemini-omni-flash and gemini-omni-pro based on the Veo 3.1 pattern.
Today: four separate calls across two vendors.
# TODAY: four API calls, two vendors, four error taxonomies
from openai import OpenAI
from google import genai
openai = OpenAI()
google = genai.Client()
transcript = openai.audio.transcriptions.create(model="whisper-1", file=audio)
vision = openai.chat.completions.create(model="gpt-4o", messages=[
{"role": "user", "content": [{"type": "image_url", "image_url": image_url}]}])
caption = openai.chat.completions.create(model="gpt-4o", messages=[
{"role": "user", "content": f"Caption for {vision.choices[0].message.content}"}])
video = google.models.generate_videos(model="veo-3.1", prompt=caption.choices[0].message.content)Tomorrow (projected): one call to Omni.
# PROJECTED: single Omni call (pseudocode — real SDK lands with the API)
from google import genai
client = genai.Client()
response = client.models.generate_content(
model="gemini-omni-flash", # or "gemini-omni-pro" for longer clips
contents=[text_prompt, image, audio, video_reference],
config={"output_modality": "video", "duration_seconds": 10}
)When the API drops, the migration is mostly deleting code. Watch for the model name on launch day, and pin it to a constant so you can flip between omni-flash and omni-pro per workload without touching call sites. The current Python SDK pattern for Gemini already handles multimodal contents, so the shape above slots in cleanly.
How Techsy Thinks About Omni Migration
When we evaluate any model-stack swap for clients, we ask three questions: what's the latency budget, does the new model cover the modalities your workload actually uses, and is there a vendor-consolidation upside worth the rebuild cost? Two of those usually answer themselves; the third is where most migrations get talked into existing for the wrong reason.
We're not recommending clients rebuild on Omni today. The API isn't out, the prices aren't published, and a projection range is a planning tool, not a green light to ship a refactor. What we'd do right now is wire the gateway layer that lets you A/B Omni against your current stack the moment the API drops. Cache the multimodal calls aggressively (image + audio inputs are deterministic enough to hit cache often), and keep a feature flag for the model name.
Building an AI stack that needs to absorb model launches without a rewrite every six weeks? Get a free consultation and we'll pressure-test where Omni fits (or doesn't).
Frequently Asked Questions
How much does Gemini Omni API cost?
As of May 19, 2026, Gemini Omni API pricing has not been published. Google's pricing page lists every Gemini model except Omni. Anchored to Veo 3.1 and Gemini 3.5 Flash, projected rates land at $1.50–$2.50 per 1M input tokens and $0.20–$0.60 per second of video output. Numbers are projections, not confirmed rates.
When will the Gemini Omni API launch?
Google said the API would roll out "in the coming weeks" at I/O 2026 on May 19. Vertex AI access typically ships first for new Gemini models, with AI Studio pay-as-you-go following a few weeks later. The official launch post is the canonical timeline source. We'll update this post the day rates publish.
Is Gemini Omni cheaper than GPT-4o?
It depends on the workload. For video output, Omni is the only Big-3 option that ships generation in one call, since GPT-4o doesn't output video. For text-only work, GPT-4o-mini at $0.15/$0.60 per 1M tokens beats projected Omni rates by roughly 3×. Pick the model that matches the modalities your pipeline actually produces.
What does Gemini Omni replace in my stack?
For multimodal video workflows, Omni replaces four separate API calls: Whisper for transcription, GPT-4o Vision for image understanding, GPT-4o for text generation, and Veo 3.1 for video synthesis. The biggest win is usually losing three vendor SDKs, three SLAs, and three error taxonomies, not raw dollar savings. See the stack-collapse worksheet above for per-clip math.
Can I use Gemini Omni via the API today?
No. As of May 19, 2026, API access is not yet available. Consumer access is live via Gemini app tiers: AI Plus at $20/mo, AI Pro at $30/mo, and AI Ultra at $100/mo, per Google's AI subscriptions blog. API rollout via Vertex AI is "coming weeks" per Google's own wording.
Does Gemini Omni support fine-tuning?
Not announced as of launch day. Veo 3.1 doesn't support fine-tuning either, and Omni is built on the same video foundation, so assume no fine-tuning at launch. Google has historically added fine-tuning to multimodal models a few months after general availability, so a Q3 or Q4 2026 timeline is plausible but unconfirmed.
How does Gemini Omni billing work?
Projected to follow Google's standard pattern: per-million-token rates for input (text, image, audio, video frames) plus per-second billing for video output. Batch (typically ~50% discount), Flex, and Priority tiers are likely available like other Gemini models. Context-tier pricing (2× rate over 200k tokens) probably applies given video frames push token counts fast.
What's the difference between Omni Flash and Omni Pro?
Omni Flash is the faster, lower-cost variant capped at 10-second clips. Omni Pro runs longer clips with higher fidelity at higher projected cost. Both cover the same modality matrix (text, image, audio, video in; video out today). Google calls the 10-second Flash limit a "deployment decision, not a model constraint," so expect it to extend.
Will Gemini Omni replace GPT-4o for my Instagram automation?
It depends on what you generate. For Reels (video output): yes, eventually, since Omni is the only single-call option for 30-second video synthesis. For static posts with AI captions and hashtags: probably not. GPT-4o-mini still beats projected Omni rates by ~3× on cost, and the n8n + GPT-4o-mini chain is already battle-tested in production.
Wrapping Up
Three things to take away:
- Gemini Omni API pricing is not published as of May 19, 2026. The pricing page row is blank, consumer tiers are live at $20–$100/mo, and API access via Vertex AI is "coming weeks."
- Projection range: $1.50–$2.50 per 1M input tokens and $0.20–$0.60 per second of video output, anchored to Veo 3.1 and Gemini 3.5 Flash. Audio input projected at $3–$5 per 1M tokens. All numbers are projections, not facts.
- The stack-collapse saving isn't always dollars. It's losing 3 vendors, 3 SDKs, and 3 SLAs in a single multimodal call. Plan the migration around vendor consolidation and latency, not raw $/clip.
Last verified: 2026-05-19 (update SLA: 24 hours after Google publishes rates). I'll update this post the day Google publishes API rates. Bookmark it.
Techsy's editorial team has shipped multimodal pipelines on GPT-4o + Veo 3.1 + Whisper since 2024. We update this post when Google publishes Omni rates.