Forgot your password?
typodupeerror
AI Google

Google's Gemini 3.7 Flash Targets Coding and Agents With a 50% Price Cut 40

Google has released Gemini 3.7 Flash just three weeks after 3.6 Flash, focusing on better coding, agentic workflows, and enterprise automation while temporarily cutting API prices in half through the end of 2026. VentureBeat reports: For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens.

Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google's claimed reductions in retries and manual oversight translate into lower total operating costs. The launch also underscores Google's rapid iteration on its Flash line while its next flagship Pro model remains absent. [...]

Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity. [...] Google's benchmarks show a large generational improvement in several software engineering tests. [...] Google's own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.
This discussion has been archived. No new comments can be posted.

Google's Gemini 3.7 Flash Targets Coding and Agents With a 50% Price Cut

Comments Filter:
  • by barc0001 ( 173002 ) on Thursday August 13, 2026 @06:18PM (#66288061)

    Be interesting to see how many of these companies bite the dust if the bigger players with deeper pockets like Google keep access pricing down longer to drive adoption and build their market share. I know many companies are already freaking out about some of the jumps in pricing that still don't pay the 'real' cost of using AI and have already aggressively cut back so the industry as a whole is in for a rough time not properly factoring in exactly how miserly their potential customer base is, but this wrinkle is going to make it way worse for some.

    • by allo ( 1728082 )

      Google cuts the price, because more powerful Chinese models were cheaper. When a Chinese company releases a good model for free, you instantly have a lot of companies (many American) that host it. As they don't have to pay for training they can operate it for hosting price plus some profit margin and undercut everyone else.

    • Be interesting to see how many of these companies bite the dust if the bigger players with deeper pockets like Google keep access pricing down longer to drive adoption and build their market share.

      Gemini's new discounted Flash pricing is still ~3x the price of GPT Luna, and that was released a week or two ago.

      Everyone's hemorrhaging money here, but Google feeling they can't match or undercut isn't a great sign for them. And Gemini is flagging in intelligence too, which would be a strong reason to underprice if they could. Gemini Flash used to be your cheapest option. I'd say structural advantage still puts the odds in Google's favor in the long run, but has to be a worse situation than they ever expe

  • Or at least heavily subsidised.
    • I think this is plan B. Google doesn't want to play the discount game; they want to be on the frontier. But they've lost some top talent and maybe they're struggling to keep up. Going value is a good backup plan.
  • At hyperscale data center rates (approx. $0.06 â" $0.08 per kWh in North America/global hubs), the raw power cost to generate 1 million tokens is only $0.008 to $0.032 (less than 3 cents).

    • by allo ( 1728082 )

      The other question is, how efficient is the hosting? In case of Gemini it is highly optimized. They split for example different tasks in the pipeline over different servers so a single server can be optimized to do a single task highly parallel.

      Measuring the environmental impact of delivering AI at Google Scale: https://arxiv.org/abs/2508.157... [arxiv.org]

  • Let's get you hooked at prices you can afford, then hope you stay when it becomes unaffordable.

  • by Tony Isaac ( 1301187 ) on Friday August 14, 2026 @12:02AM (#66288337) Homepage

    They're more like airline miles, which are defined to be worth whatever each airline wants them to be worth.

    Each AI company defines tokens somewhat differently, so it's not possible to compare costs-per-token between the different companies' products.

    Right now, Copilot seems to be the highest price of the big AI coding agents. It's *really* easy to churn through $20 in tokens in an hour. Claude, on the other hand, performs a similar amount of work for $20 a month.

    • Are you tweaking your reasoning level? That's becoming more important lately when managing your token burn.

      I'm using Deepseek Pro until a better quality per cost option comes along, but I burned through about 100 million tokens per day ($8.70 at current prices) set at Maximum reasoning. You're basically telling it to torch tokens at ludicrous overkill, but the output is was exactly what I needed for the task. I finished a massive load of work in a few days that I had been stalled on for months because my
      • On Gihub Copilot, I use "Auto" most of the time. It's supposed to pick the model best suited for each request, and provides a 10% discount. It also does not use the latest, greatest, most expensive models using this setting. So no, I don't think I triggered the high costs by using inappropriately powerful or expensive models.

        • Also, another consideration...input token cost.

          I just made this mistake. I was in a long-running programming session with a lot of context, and I asked an unrelated question, "how to slob a knob" or something like that.

          I constantly check my balance and noticed I burned 200k tokens with a simple question.

          Then it dawned on me. It included every scrap of information in the session with the request, as it does, juicing me for a mountain of unnecessary input tokens.

          So, lesson learned. For each unrela
    • Each AI company defines tokens somewhat differently, so it's not possible to compare costs-per-token between the different companies' products.
      Perhaps how tokens are used in "calculations".

      Not how tokens are generated from input.

  • And as soon as you are, expect to get fleeced.

  • As the stories come out about companies using up their tokens, and there's still *no* path to profitability.

  • What is a million output tokens comparable to? How many lines of average-length code, over how many attempts/corrections? Just to give us a better understanding of what one would be paying 3.75 for, and if it's worth it.

  • I'd rather spend an entire day writing code, learning something, keeping my brain engaged, in the state of flow, and actually enjoying it than trying to correct a mediocre agent all day, getting frustrated and burnt out, and still having to pay for it.

"You show me an American who can keep his mouth shut and I'll eat him." -- Newspaperman from Frank Capra's _Meet_John_Doe_

Working...