I Asked Claude to Build a Crystal Ball for Flight Prices
The Eras Tour isn't one of the things that move a plane ticket's price. I checked.
I wanted to predict them. What I got was something better, and a lot more honest about what these tools are actually for.
A handful of things actually move the price of a plane ticket. A lot of things you’d swear move it don’t. The Eras Tour, I can now report, isn’t one of them.
Yes, I fly Delta constantly. (To anyone who knows me: you know I have a borderline toxic relationship with it.) Domestic travel all over most months, London a couple times a year, and somewhere in all that booking I got tired of doing the math in my head: is this a good price, is it going to drop, should I just hit buy already. So I did what anyone with an AI and a little too much curiosity does. I asked it to build me a model that predicts flight prices.
It did not go how I planned. That’s the fun part.
The fifteen-year graveyard
Here’s the thing I didn’t fully believe when I started… people have tried and failed at this and I was not going to fix the problem with the best AI models 2026 had to offer.
People have been trying to predict flight prices for fifteen years, at every scale, with budgets I will never have. Farecast, the original, got good enough that Microsoft reportedly paid $115 million for it back in 2008. Then quietly shut it down in 2014. FLYR Labs started out selling travelers a way to lock in a fare, then in 2019 decided the consumer side wasn’t the business, pivoted to selling its forecasting engine to the airlines instead, and went on to raise hundreds of millions doing exactly that. Hopper will tell you it’s “95% accurate,” a number that lives entirely in affiliate blog posts and is backstopped by a refund policy, not a methodology. Even Kayak, the most honest of the bunch, only claims around 80% three weeks out, and past a couple of months it’s worse than a coin flip.
And Google Flights, the one you actually use? That “prices are low right now” bar isn’t predicting anything. Google says so itself. It’s just comparing today against the recent past.
There’s a pattern in the graveyard worth noticing: nobody who chased this stayed in the business of predicting fares for travelers. They either walked upstream and sold the engine to the airlines, or sideways into simply organizing the search. And the people who obsess over this for sport, the Reddit fare-hunters and the points bloggers, land in exactly the same place the academics do: a plane ticket is basically a random walk with a fuel-shock tail. The crystal ball itself never paid, for anyone.
Claude tried to warn me. Fifteen years, billions of price points, and the answer is still nobody really knows. Naturally, I rolled up my sleeves and tried anyways. I had Fable, GPT 5.6 Sol, Perplexity… I was armed (or so I thought).
Finding zero, in my own data
The plan was gloriously overbuilt. I had my AI pull every variable I could dream up and throw them all at the question. Jet fuel prices. Geopolitical risk indexes. TSA checkpoint volume. How many seats were left on the plane. Storms. Consumer confidence. And, because at some point the whole thing tips into mad science, the price of milk and eggs.
This wasn’t a back-of-the-napkin thing. I pulled a booking-curve panel out of 82 million real fares, sampled five thousand flights across 234 routes, then checked it against 39 of the routes I actually fly, thousands of prices tracked from six months out to about four weeks before departure. Twenty-eight variables, from a much longer list. I wasn’t playing around. Who said a year of PhD statistics would never get used?
I was sure the exotic stuff would be the secret. It felt like the secret. Surely war and fuel and demand were the edge the big players were missing.
They were not. Almost none of it held. The signal everyone in the field loves most, how many seats are left on the plane, barely moved the needle. Fuel, geopolitics, TSA lines: mostly decoration. (Storms left a faint fingerprint, actually, but system-wide and a few days later, never on the one flight I was trying to book.)
Then there’s the one everybody assumes. A big event in town, surely, sends flights through the roof. The Eras Tour boosted the economy by a reported $5 billion; surely that — or the Super Bowl (perhaps to catch her now husband) — moves the needle. It doesn’t. When a huge event lands in a city, hotel prices go feral (up fifty, a hundred, several hundred percent) and airfare barely blinks. The hotels absorb the mania; the flights just shrug. So no, the Eras Tour does not move airfare. It only moved me, and my hotel bill.
And then there was milk. Milk almost worked. In one test it came back borderline statistically significant at predicting airfare, the kind of result that makes you sit up in your chair. For about ten minutes, I thought I’d found something genuinely stupid and wonderful.
Here’s why it’s a beautiful lie. Milk and jet fuel ride the same cost curve; fuel, transport, and inflation all push on both at once. So milk and airfares drift together, not because milk tells you anything about your flight, but because they’re both quietly holding hands with the same forces underneath. Line it up as a leading indicator, ask whether the price of milk today tells you anything about your airfare next quarter, and the effect vanishes completely. Eggs never even pretended. I heard my dissertation chair Fred Galloway in the back of my head: correlation does not mean causation.
The red-team twist
I didn’t trust my own results. When you want a model to work, you’re the worst person to check it. So I turned the AI against its own work. Two of them, actually, set loose as adversaries with one job. Tear this apart. Assume I fooled myself. Find where. I built a custom skill just for this task.
My favorite piece of the whole build was the part that tried to read the weather of the market: fuel spikes, demand swings, the general conditions a fare is sold into. It felt sophisticated. One of the adversaries took about a paragraph to show it was, essentially, the calendar. The apparent intelligence was just booking season, dressed up as insight.
The single most useful thing my model did was tell me whether today’s price was high or low for what the flight actually is. Which is, more or less, what that little Google Flights bar already does. I had rebuilt Google’s modest scope, quantified it, and confirmed it the hard way.
What actually had to be true
Somewhere in all that wreckage I realized I’d been running the same four tests on everything and had never once written them down. Does the thing actually move fares. Does it move them at the grain I care about, meaning this flight in the next two weeks and not the whole market over a year. Does it tell me something before it happens instead of at the same time. And does any of it beat just booking the ticket.
Almost nothing cleared all four. The two that came closest each failed on a different one, and honestly the failures are the more interesting half.
Fuel: passed one test, failed another
Fuel is the one that passed the first test and I still couldn’t use it.
It genuinely moves airfares. That part isn’t in question, it holds up across the whole market over months and the pass-through is real. Test one, cleanly. Test two is where it dies, because moving the whole market over months tells me nothing about which flight next Tuesday is the one to book. Right signal, wrong grain.
By now some of you are thinking: fine, but if fuel moves airfares, why not just predict fuel? Oil is harder to forecast than airfare. It’s the most-watched commodity on earth, and its biggest moves are geopolitical shocks nobody sees coming. Predicting fuel just moves the crystal ball up one shelf and hopes you don’t look.
So it didn’t go in the model. It went where it belongs, which is a separate read on whether the whole market is running hot. What fuel can tell you isn’t when to book, it’s whether to stop waiting: when fares are high because of a shock that takes months to clear, you quit holding out for a “normal” that isn’t coming back this season.
The call
Which brings me to the fourth test, the only one that actually pays. Does any of this beat just booking the ticket.
The model does, a little. Quietly watching a route and pouncing on a real dip does considerably better. When I ran the numbers it came out around $20 a booking for the watching against three to five for the prediction, which is not the ratio I expected going in and is, I think, the whole finding. I wanted a crystal ball. What I got was a good pair of eyes.
The part I keep coming back to isn’t about flights though. The most valuable thing the AI did all month was help me kill the piece of the model I was proudest of. Left to my own devices I’d have shipped that market-weather read, told you it was clever, and quietly believed it myself. Its job was to argue with me and to refuse to flatter. AI isn’t going to replace the thinking, its real gift is making the thinking better. It didn’t hand me the answer, it made sure I didn’t fool myself into a worse one. There’s a bigger piece in that last thought and I’ll write it. Soon. For now, back to the flights.
So how do I actually do the math now?
Less magically than I hoped, and better than I did before.
I did build a model. It’s just a humble one, and it knows it. It looks at four plain things: the price itself, how far out I’m booking, which route it is, and how that price compares to the route’s own normal fare for that point in the booking window. From those, it estimates a single number: the chance the fare drops by a meaningful amount in the next two weeks, with honest error bars, and it flags when it isn’t sure. That’s the whole engine. The exotic stuff I was so sure about (the seat counts, the storm flags, the geopolitics) never made it in. Four boring inputs did all the work.
IN — the price · how far out I’m booking · which route · how it compares to the route’s own normal fare
↓
OUT — the chance this fare drops by a meaningful amount within 14 days
(For the stats nerds still with me: it’s a gradient-boosted classifier with isotonic calibration, cross-validated AUC of 0.745 against a coin-flip-ish 53% base rate. In a separate logistic fit, the fair-value term lands at β ≈ 4.5. Strip it down to that one feature and you keep most of the lift, which is the rigorous way of admitting the model is mostly one variable and a lot of restraint about the rest.)
Line up two fares, one that’s about to drop and one that isn’t, and the model picks the right one about three times out of four.
The fine print: on the thin little route I fly most, the real edge is a few dollars a booking, not a fortune. The model has also never once seen a true last-minute scramble, which is exactly the moment you’re most desperate for it.
And the model isn’t even the part I lean on most. The heavy lifting is the watching. A little system tracks the routes I fly, knows what a genuinely good price looks like from the actual history, leaves me alone when nothing’s happening, and pings me the moment a fare drops into “book it” territory. And for the part I care about most, it works out which way of booking actually earns the miles and status I’m chasing, the kind of arithmetic no human should be doing at 11pm.
Not a crystal ball. A model that knows what it doesn’t know, a patient watcher, and a nudge at the right moment.
Which is what my toxic relationship with Delta needed all along — not something to tell me the future, just something to keep me from doing something dumb, and to notice the good moment when it comes.
Try it yourself
If you want to try the thinking without the infrastructure, here’s the prompt. Hand it to whatever AI you use:
You are helping me decide how to book a flight. Inputs: origin(s) (including plausible alternates I could drive to), destination, dates (fixed, or a flex window of ±N days), cabin, and my airline/alliance constraint. Do not price a single itinerary. Enumerate the realistic booking configurations: round-trip vs sum-of-two-one-ways vs separate tickets through a nearby gateway, across each date in my flex window and each viable cabin, filtered to my alliance. Price each configuration live on Google Flights (or equivalent).
For every configuration, anchor the current fare against that route's own recent price history — Google Flights' price-history graph and its "prices are currently low/typical/high" insight: where does today's price sit in the last ~60 days' range (near its floor, typical, or high)? Then compare configurations at true trip cost, not sticker price: add airport parking or positioning costs for alternate origins, a required destination hotel night whenever a connecting international outbound obliges landing a day early (a nonstop may land day-of), and flag separate-ticket risk (no misconnect protection, bags re-checked). If a loyalty program matters, annotate each configuration's status-earning — noting that fare-based and distance-based earning diverge on codeshares, so the marketing carrier choice on an identical flight can double the status value.
For the recommendation, do not forecast whether fares will rise or fall — short-horizon fare direction is effectively random; treat any confident "prices will drop" prediction as noise. Give a conditional read instead: if the current fare sits at or near the route's historical floor, or departure is within ~3–4 weeks (when fares mostly rise), recommend booking now; if the fare is high against its own history and departure is far out, recommend waiting with a concrete trigger — name the target price (the route's typical or floor level) at which I should book immediately without re-deliberating. Present the configurations as a ranked table with one clear pick and the reasoning, never a bare score.
One caveat if you take that prompt and run it: the judgment transfers, the math doesn’t. When my system calls a fare a steal, that’s not an eyeball of today’s Google Flights graph, it’s a statistical read against more than 14,000 price observations my server has been quietly collecting on my actual routes for months. It knows LA to Heathrow actually bottoms out around $323 one-way (not the $374 the graph calls “typical,” a distinction that took several hundred observations to see). It scores today’s fare against that distribution, flags a genuine outlier, and runs a small calibrated model estimating the odds of a meaningful drop in the next two weeks (AUC 0.745, which is modest, and I won’t pretend it’s a crystal ball; that was sort of the point).
And the prompt is one-shot where the real thing never stops looking. Your AI reads the graph once and hands you a trigger price; mine re-prices the route twice a day, re-arms after a dip recovers, and pings me exactly once when a decision point actually arrives. None of that is producible in the moment, and no prompt fixes it. The transferable part is the thinking: enumerate the configurations, anchor against the route’s own history, refuse to forecast. The rest of the edge is admittedly boring, accumulated data and a computer that keeps paying attention long after you’ve closed the tab.
Building something with these tools — and getting kept honest by one? Send a note.