
Earlier this month, we ran our first AI Summer School, a series of webinars where restaurant operators demoed real tools and workflows they’ve built with AI. Phil Smith, Director of Marketing at Upstream Hospitality, joined me for one of those sessions to walk through how he uses Bikky AI in his day-to-day work.
Phil is about as heavy an AI user as you’ll find in the industry. He’s using Claude daily, and in an earlier Summer School session he walked us through how he uses it to build and execute a paid ad campaign for a new store opening. This time, instead of a single project, we walked through how he uses AI in his day-to-day decision making.
The session reinforced the current state of AI: the real barrier isn't how sophisticated your questions are, it's how good the data foundation underneath the model is.
Phil walked us through a special Upstream has run for years: a mussels and beer combo on Monday nights, a nod to the mussels that were the signature dish when their first location opened back in 2011. Mondays are generally their slow day, so Phil wanted to refresh the special, and do something on-brand for football season. So he asked Bikky AI a simple question: “is this working, and if not, what else could we run instead?”
The first answer surprised him. Even though the mussels special wasn't moving a lot of volume, the checks it appeared on were carrying real revenue. Cutting it wasn't as simple as looking at the item count and deciding it needed to be replaced. As Phil put it, the real question wasn't whether people were buying mussels, it was what happened to the whole check when they did.
That's an easy thing to miss if you're just looking at a POS export or the PMIX. It doesn't tell you who bought it, what else was on their check, or whether they came back.
In the end, Phil didn't kill the special because it was underperforming in isolation. Instead, he replaced it due to something else he discovered in his conversation with his data. A smash burger they had launched back in January was already selling 50% more on Mondays than other days, and had a 90-day retention rate of over 27%. That insight became "Get Smashed," a burger and beer combo built for football season, incremental to check, and with broader appeal than a dish whose signature status had faded as the brand had grown.
As Phil worked through what to do about the special, he didn’t have to toggle between reports on mussel sales, burger performance, and guest retention. He stayed in a single thread, pulling in new metrics as his thinking evolved, with confidence that every response was accurate.
The second example was smaller in scale but made a similar point. Upstream launched a new chicken tender item in January with four size options: 8, 10, 12, and 20 pieces. When it came time to build the new menu, Phil wanted to check his instincts against the data.
He asked Bikky AI for the differences in checks by tender size. He found that those ordering 10 pieces vs. 12 were functionally the same guest - with little difference in check, daypart, and occasion. This meant that 12-piece actually carried a worse food cost than the 10, since guests were getting more food for essentially the same price. Keeping both sizes wasn’t really giving guests more choice; it was just shuffling guests around the menu with no real benefit. Bikky AI suggested he test cutting the item to see if those guests would shift to the 10-piece, as expected.
But what Phil did next was the part I found most interesting. He took the same question to Claude, independently, and asked it to challenge the recommendation. He wanted a second opinion and pushed it to find a reason the answer should be different.
Claude landed in essentially the same place, but with a slightly more aggressive prescription: cut the 12 piece outright.
Not because Claude and Bikky AI are the same model, but because they were reasoning from the same underlying picture of the business.
Phil was pretty direct about how he thinks about trust in the data:
“Bikky is the source of truth. What I’m asking Claude for is to be another objective reasoning model, to go through everything logically with the numbers that came from Bikky.”
That distinction matters. The value isn’t necessarily in having one magic AI model. It’s in giving the models a reliable picture of the business to reason over.
The last example started with an almost throwaway question. Upstream was running a giveaway for a free year at Taproom and needed a dollar value for the prize.
The old way to answer that might be to take an average check, multiply it by an assumed number of visits, and call it a day. Instead, Phil asked Bikky AI what a year at Taproom is actually worth.
The answer was $1,200, but what made it compelling wasn't the number, it was the reasoning behind it. Bikky AI started by pointing out something obvious in hindsight: an average across every guest gets dragged down by one-time visitors who never come back, and those guests were never who this promotion was built for. So it narrowed the analysis to the people most likely to actually treat 'a year of Taproom' as a year of real visits, guests ordering 24 or more times annually. Among that group, the median annual spend was $1,007 and the average was $1,322. $1,200 landed right in the middle, and came out to a clean $100 a month. It even flagged a more conservative option, $1,000, which tracked almost exactly with the median spend of the highest-frequency guests.
Phil hadn't asked for any of that. He hadn't specified which guests to look at or how to handle outliers. The system worked out on its own that a frequent Taproom guest was a far more relevant comparison for this question than an average one, and it showed enough of its work that Phil could trust the number without redoing the math himself.
That was the pattern I kept seeing throughout the session. These are the kinds of questions operators ask all the time. Should we keep this special? Do we need this menu size? What’s a year's worth of visits actually worth?
What used to make them hard was the hours of spreadsheet work standing between the question and a trustworthy answer. Exporting, cleaning, merging - anyone who’s dug through this process knows that the time consuming part isn’t plugging in the formulas for the analysis, it’s prepping the data for analysis in the first place.
You have to know that the different ways an item gets entered in the POS actually refer to the same thing. You have to know that the same guest ordering through different channels is actually the same guest. You have to understand what the numbers mean in the context of that particular restaurant. And you have to be able to look across those pieces at the same time.
That’s what I found most interesting about what Phil shared. The AI wasn’t replacing his judgment. It was making it dramatically easier to get to the information he needed to exercise that judgment.
A lot of the conversation about AI and restaurants right now is about people starting to get access to tools like Claude and ChatGPT. That’s an important first step, and we’re already seeing operators roll out these tools across their teams. But once everyone has access, a different question starts to matter: how do you make sure there’s consensus and trust in what the models are running on?
The models will keep getting better. They’ll reason faster, handle more complex problems, and become cheaper and easier to use. But they still reason over whatever you give them. If the underlying picture of the business is incomplete, inconsistent, or missing important context, AI just gets to the same wrong conclusion a human would, just much faster.
But if the foundation is right, something much more interesting happens. AI can start becoming the first step for working through questions that actually drive your business.
As intelligence becomes more abundant, I think that foundation is going to matter more, not less.