Two years ago the entire AI conversation was about Token Maxxing, chasing the newest model and pushing as many tokens through the system as possible. I have written about this before. What's happening right now is the next chapter, and it looks nothing like the last one.
The question is no longer which model is biggest. It's which model earns its token count.
Kimi K3 landed this month and it is worth watching. It runs $3 input and $15 output per million tokens, roughly 40 percent cheaper than Claude Opus 4.8. It is not a clean sweep, Claude still leads on the harder production coding work, but K3 wins outright on frontend generation and several agentic tests. Cost and capability used to move in the same direction. They are starting to split.
OpenAI is telling the same story from a different angle. GPT-5.6 now ships in three tiers instead of one: Sol for frontier reasoning, Terra as the balanced daily workhorse, Luna for fast, high-volume tasks. Spinning up a flagship model to reformat a script is like flying a commercial jet to the grocery store. The tier system exists because not every prompt deserves the same horsepower.
Here is where it gets genuinely interesting. Cursor just launched something called Cursor Router, live for a few days now. It is not a new model. It is a decision layer that reads every request and sends it to whichever model actually fits the job, claiming frontier quality at 60 percent lower cost. Auto mode inside Cursor runs on it by default now.
That is the industry admitting, that intelligent routing is infrastructure now.
It is also worth asking who owns that decision layer. Cursor is mid acquisition by SpaceX in a deal built to strengthen Grok. Whoever controls the router controls where the token spend actually flows, and that is a bigger strategic asset than the market seems to be pricing in right now. There is a deeper flywheel story underneath this, Tesla's data, Starlink, SpaceX's valuation. That's a post on its own.
#AI #LLMs #TokenEfficiency #Cursor #OpenAI #KimiK3