Google has reportedly hit Meta with a usage cap on its Gemini AI model after the social media giant blew through its computing allocation. That is not a great look for two of the best-funded companies on the planet.
According to the Financial Times, Meta had been using Gemini for customer service chatbots, advertising tools, coding assistance, harmful content takedowns, and scam detection. The company chose Gemini because it outperformed its own Llama open-source models on those tasks. Meta also uses Anthropic’s Claude for similar workloads.
Google flagged the capacity issue back in March, forcing Meta to tell employees to use tokens more efficiently. The reality is stark: even companies building their own massive AI infrastructure cannot get enough compute to keep up with demand.
Meta has pledged $600 billion toward US data center construction over the next two years. It does not run its own cloud business, making it dependent on partners like Google for capacity. Meanwhile Google itself recently agreed to pay SpaceX $920 million per month for extra data center capacity to handle Gemini Enterprise workloads.
The ripple effects are spreading. Token prices have surged. Some companies are pulling back on AI usage entirely, telling staff to stop using AI for anything non-essential. The irony: the companies building these models are also struggling to afford running them at scale.
We are in a phase where AI capabilities are advancing faster than the infrastructure can support them. Until that gap closes, expect more caps, more trade-offs, and more hard choices about who gets to use the models and how much.
