Google's Gemini 3.7 Flash Targets Coding and Agents With a 50% Price Cut 40
Google has released Gemini 3.7 Flash just three weeks after 3.6 Flash, focusing on better coding, agentic workflows, and enterprise automation while temporarily cutting API prices in half through the end of 2026. VentureBeat reports: For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens.
Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google's claimed reductions in retries and manual oversight translate into lower total operating costs. The launch also underscores Google's rapid iteration on its Flash line while its next flagship Pro model remains absent. [...]
Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity. [...] Google's benchmarks show a large generational improvement in several software engineering tests. [...] Google's own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.
Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google's claimed reductions in retries and manual oversight translate into lower total operating costs. The launch also underscores Google's rapid iteration on its Flash line while its next flagship Pro model remains absent. [...]
Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity. [...] Google's benchmarks show a large generational improvement in several software engineering tests. [...] Google's own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.
Begun, the AI price wars have... (Score:4, Interesting)
Be interesting to see how many of these companies bite the dust if the bigger players with deeper pockets like Google keep access pricing down longer to drive adoption and build their market share. I know many companies are already freaking out about some of the jumps in pricing that still don't pay the 'real' cost of using AI and have already aggressively cut back so the industry as a whole is in for a rough time not properly factoring in exactly how miserly their potential customer base is, but this wrinkle is going to make it way worse for some.
Re:Begun, the AI price wars have... (Score:5, Interesting)
You'd be astonished at how many companies will be absolutely fine with you driving that Yugo as long as it's cheaper.
The company I work at has top down shifted which model is tied into all the devs IDEs 4 times in 6 weeks, always seeking cheaper usage even at the cost of suitability. And the company is a multinational with 25K employees worldwide, not some little shop... If my company's doing it, others are doing it too.
Re: Begun, the AI price wars have... (Score:3, Interesting)
Re: Begun, the AI price wars have... (Score:5, Insightful)
> We also have a small cluster of 7900xtx for local inference when you go over your allowance
That's actually one of the conversations I've been hearing at work is that 'free' LLMs we can self host at one of our datacenters is the long term plan once they're Good Enough(tm). I think a lot of companies are planning the same. Medium to big corporations with large tech departments remember vividly how Broadcom screwed them on VMWare pricing and they're not going to get caught in the same bear trap with AI. They're already planning to be mostly to fully self sufficient - which again is going to cause a lot of those "AI will be worth XXXXXXXX" predictions to be completely and utterly wrong and kill a lot of those companies.
Re: (Score:1)
As a local AI enthusiast I can tell you that it's not practical even with a large budget. For one thing the hardware is obsolete as soon as you buy it. In a few years that $30k compute card is worth $300. Are you or your company willing to completely replace a multimillion dollar infrastructure every 2 or 3 years? It's much cheaper to pay the fees and always have the latest hardware.
With that said, even "old" AI is useful for some tasks but probably not what you'll want to use it for. The field is moving so
Re: (Score:2)
Also, the big AI companies get a number of advantages, such as large batch sizes and being able to balance on-peak and off-peak workloads better.
Re: (Score:2)
In a few years that $30k compute card is worth $300. Are you or your company willing to completely replace a multimillion dollar infrastructure every 2 or 3 years?
At the moment all of the big AI providers building data centers are paying 10x the normal price for GPUs and RAM which is why companies like NVidia and Micron are making out like bandits. Keep in mind the AI providers have to recoup that spending at some point. Of course the rest of us are also paying 10x normal prices at the moment if we attempt to build our own servers.
NVidia and Micron (TSMC, Samsung etc) must have been massively ramping up production capacity to meet the demand.
Not only that but every
Re: (Score:2)
I don't think you are going to see that flood of cheap hardware.
The data center kit is not the relatively well behaved mostly PC compatible parts of the late 90s, 2000s, or even 2010s. This stuff is super dense with special power and interconnect requirements normies are just not going to do in their homes.
I mean sure some nerds with separate electrical panel and equipment racks in the basements will (they always do) but the effort level here is going to look more like firing up an old VAX than some 2U x86
Re: (Score:1)
At the moment all of the big AI providers building data centers are paying 10x the normal price for GPUs and RAM which is why companies like NVidia and Micron are making out like bandits. Keep in mind the AI providers have to recoup that spending at some point. Of course the rest of us are also paying 10x normal prices at the moment if we attempt to build our own servers.
NVidia and Micron (TSMC, Samsung etc) must have been massively ramping up production capacity to meet the demand.
Not only that but every year Moore's Law rolls on and GPUs and RAM keep getting faster and cheaper. Obviously this hasn't been happening for the last few years because demand is so high that even GPUs such as the 3090 which is 6 years old and should be selling for a tenth of its original RRP is still selling SECOND HAND at the same price. This is not normal.
My question is what happens when these huge buildouts eventually start to slowdown?
Separate from demand AI has also become a convenient excuse to artificially restrict supply with DRAM now having higher margins than HBM.
Is it ironic that AI companies going all out on data centers is the only thing keeping cheap powerful AI out of the hands of the masses? Once they stop building data centres hardware prices will normalize and their valuations will tank.
The thing I keep waiting for is for someone to create an AI centric memory controller that uses a shit ton of dirt cheap DRAM in parallel... all it has to do is block transfer into specific regions of a processors SRAM to keep it fed. Would be able to cut out huge amounts of complexity and synchronization/power constraints vs general purpose GPUs. Perhaps go a bit furthe
Re: (Score:2)
There are some AI applications that need to be hosted on-prem due to privacy reasons and are not even connected to the Internet. One example is an internal tool that gets a feed of logs from all other applications and performs real-time analysis and generates alerts as well as suggests root-cause and potential fix, etc.
Re: (Score:2)
As a local AI enthusiast I can tell you that it's not practical even with a large budget. For one thing the hardware is obsolete as soon as you buy it. In a few years that $30k compute card is worth $300.
This would be an awesome problem to have. My local AI enthusiast hardware is three years old and currently worth 4x what I paid for it at the time. At least when shit becomes worthless you can afford to buy better shit and get more value out of it.
With that said, even "old" AI is useful for some tasks but probably not what you'll want to use it for. The field is moving so insanely fast that newer models vastly surpass previous models and you want to be top top of that because the improvements are huge.
Personally don't see hardware requirements moving all that much. Much of it as I predicted three years ago is in large sparse models with limited numbers of active parameters. I can run everything I would want to run except K3 with good enough performance for
Re: (Score:2)
Medium to big corporations with large tech departments remember vividly how Broadcom screwed them on VMWare pricing and they're not going to get caught in the same bear trap with AI.
How do you reduce VMware to a "trap" when it worked fine for 20+ years? Two decades for a competitor to answer VMFS, Vmotion, and a control plan that doesn't suck.
I don't know a single soul that regrets using it in all that time, because we knew then what the competition looked like, and we sure as hell all know what it looks like now. Sorry man, but "there's nothing better" is not a trap. Things change and we're looking for alternatives and they still suck, as they always have, so it's suck vs pay more. I
Re: (Score:2)
the company I work at has top down shifted which model is tied into all the devs IDEs 4 times in 6 weeks, always seeking cheaper usage even at the cost of suitability. And the company is a multinational with 25K employees worldwide, not some little shop...
That isn't really surprising. Large business are notorious for their levels of waste but generally speaking when you have a op-ex line item that scales directly with the number of a employees in a whole department it isn't hard for the bean counters to look at that start asking simple questions like:
Does everyone really need this?
Can we get some version of this cheaper per head count?
How did our favorite productivity metrics look over the eight quarters before we deployed this vs since; how much head count
Re: (Score:2)
Is most of your daily driving "a race"?
This is a midrange model. This is competitive midrange pricing. For my midrange work I've been using GLM-5.2 - Gemini 3.7 flash is a bit cheaper and a bit better on the benchmarks, so yeah, I'll probably switch so long as the prices stay low.
Sometimes you need a high-end model (Claude Fable, GPT 5.6 Sol, Kimi K3, etc). Sometimes a low-end model is fine (GPT 5.6 Luna, DeepSeek 4 Flash 0731, etc). Usually a midrange model is the best balance.
Re: (Score:2)
You're need to get to work, get groceries, maybe go for an occasional Sunday drive. Do you buy a Porsche or a Honda?
Re: (Score:2)
Depends if I am young, single (so don't need more than one bag of groceries or so), and have enough disposable income - the Porsche, otherwise the Honda.
Re: (Score:2)
There you go.
What's your company's policy on expensing rental cars for travel? Do they cover the Porsche?
Re: (Score:2)
Google cuts the price, because more powerful Chinese models were cheaper. When a Chinese company releases a good model for free, you instantly have a lot of companies (many American) that host it. As they don't have to pay for training they can operate it for hosting price plus some profit margin and undercut everyone else.
Re: (Score:2)
Be interesting to see how many of these companies bite the dust if the bigger players with deeper pockets like Google keep access pricing down longer to drive adoption and build their market share.
Gemini's new discounted Flash pricing is still ~3x the price of GPT Luna, and that was released a week or two ago.
Everyone's hemorrhaging money here, but Google feeling they can't match or undercut isn't a great sign for them. And Gemini is flagging in intelligence too, which would be a strong reason to underprice if they could. Gemini Flash used to be your cheapest option. I'd say structural advantage still puts the odds in Google's favor in the long run, but has to be a worse situation than they ever expe
First hit is free (Score:2)
Re: First hit is free (Score:2)
Gemini, what is google's electricity cost (Score:2)
At hyperscale data center rates (approx. $0.06 â" $0.08 per kWh in North America/global hubs), the raw power cost to generate 1 million tokens is only $0.008 to $0.032 (less than 3 cents).
Re: (Score:2)
The other question is, how efficient is the hosting? In case of Gemini it is highly optimized. They split for example different tasks in the pipeline over different servers so a single server can be optimized to do a single task highly parallel.
Measuring the environmental impact of delivering AI at Google Scale: https://arxiv.org/abs/2508.157... [arxiv.org]
"That means the current discount is temporary" (Score:2)
Let's get you hooked at prices you can afford, then hope you stay when it becomes unaffordable.
Tokens are not a consistent unit (Score:4, Insightful)
They're more like airline miles, which are defined to be worth whatever each airline wants them to be worth.
Each AI company defines tokens somewhat differently, so it's not possible to compare costs-per-token between the different companies' products.
Right now, Copilot seems to be the highest price of the big AI coding agents. It's *really* easy to churn through $20 in tokens in an hour. Claude, on the other hand, performs a similar amount of work for $20 a month.
Re: (Score:2)
I'm using Deepseek Pro until a better quality per cost option comes along, but I burned through about 100 million tokens per day ($8.70 at current prices) set at Maximum reasoning. You're basically telling it to torch tokens at ludicrous overkill, but the output is was exactly what I needed for the task. I finished a massive load of work in a few days that I had been stalled on for months because my
Re: (Score:2)
On Gihub Copilot, I use "Auto" most of the time. It's supposed to pick the model best suited for each request, and provides a 10% discount. It also does not use the latest, greatest, most expensive models using this setting. So no, I don't think I triggered the high costs by using inappropriately powerful or expensive models.
Re: (Score:2)
I just made this mistake. I was in a long-running programming session with a lot of context, and I asked an unrelated question, "how to slob a knob" or something like that.
I constantly check my balance and noticed I burned 200k tokens with a simple question.
Then it dawned on me. It included every scrap of information in the session with the request, as it does, juicing me for a mountain of unnecessary input tokens.
So, lesson learned. For each unrela
Re: (Score:2)
Well, to be fair, you did ask it how to slob a knob! That would definitely cause a lot of token use! :-)
Re: (Score:2)
Each AI company defines tokens somewhat differently, so it's not possible to compare costs-per-token between the different companies' products.
Perhaps how tokens are used in "calculations".
Not how tokens are generated from input.
Re: (Score:3)
Token counts not only vary from one company to another, but from one model to another from the same company.
https://playcode.io/blog/real-... [playcode.io].
Re: (Score:2)
Thanx, good info.
Getting you hooked is free or cheap (Score:2)
And as soon as you are, expect to get fleeced.
Desperation (Score:2)
As the stories come out about companies using up their tokens, and there's still *no* path to profitability.
Illustrative Example Needed (Score:2)
What is a million output tokens comparable to? How many lines of average-length code, over how many attempts/corrections? Just to give us a better understanding of what one would be paying 3.75 for, and if it's worth it.
No, thanks. (Score:2)
I'd rather spend an entire day writing code, learning something, keeping my brain engaged, in the state of flow, and actually enjoying it than trying to correct a mediocre agent all day, getting frustrated and burnt out, and still having to pay for it.