From Gut Feel to Real Signal: Webinar Recording + Playbook
UX is shifting from gut-feel art to measurable science. David Sternberg and Eran Dror on adaptive interfaces, why simulated users are hypotheses not evidence, and where nudging turns into manipulation.

UX used to run on gut feel. A designer's instinct, a stakeholder's opinion, a coin flip dressed up as a decision. David Sternberg, Design Principal at NiaHealth and author of The Flow Equation, joined me to make the case that UX is turning into a real science, and to walk through what that actually changes about how you build.
Measurement isn't the same as science.
Webinar video
Follow us on LinkedIn to get invited to our monthly webinars →
Follow our YouTube channel for past webinar recordings →
TL;DR
- Don't just hypothesize what will work. Hypothesize why. Name the mechanism behind the behavior, and the evidence that would prove you wrong, not just the metric you expect to move.
- Simulated users generate hypotheses. They are not research. "Jane, 34, abandoned checkout" isn't evidence. Jane doesn't exist.
- As interfaces go generative, the job moves up a level. From "where does this button go" to "under what conditions should this button exist."
- The line between adaptive UX and manipulation is agency. Understanding a user's confusion to help them decide is good design. Exploiting it to steer them is not.
- Use more than one model on purpose. Living inside a single tool narrows how you think, not just what you build.
Why UX is shifting from gut-feel art to measurable science
David opens with a distinction most teams skip past: an intervention hypothesis ("simplify this flow, conversion goes up") is not the same thing as a mechanism hypothesis (why are people actually abandoning: uncertainty, cognitive resistance, a broken mental model, loss of trust). UX is already good at the first kind of question. It measures clicks, conversion, time on task, constantly. What it's weak on is a real model of the dynamics underneath words like "friction" and "flow."
His analogy is the one that sticks: people measured the movements of planets for centuries before Newton. Newton's breakthrough wasn't better measurement. It was proposing an underlying mathematical system that explained what had been observed and predicted what would happen next. That's the gap between measurement and science, and it's the gap David is trying to close in UX with a framework he calls QFIT, borrowing structure from quantum cognition and fluid dynamics to make "fuzzy" concepts like friction and confusion mathematically real instead of just descriptive.
Playbook move: before you write an A/B test brief, force a one-line mechanism hypothesis ("users abandon because X breaks their mental model of Y") and name the evidence that would falsify it, not just the metric you expect to move.
Adaptive products: apps that notice confusion and reshape themselves
"Our fixed interface is dead," is how I framed it on the call, and David didn't push back. Apps are starting to notice confusion or overwhelm and adjust in response, and AI can construct a custom interface on the fly for whatever question a user actually has. David calls generative interface design one of the largest shifts in the history of the discipline. Not a feature. A different substrate.
I used our own build as the live example. We shipped a "setup Evermuse" flow that runs as a single command inside Claude or ChatGPT, with no hard-coded UI at all, and it still has to feel like our product. That turned out to be a lot harder than it sounds. We're still not perfect at it.
Playbook move: prototype one flow as a set of adaptive rules and conditions instead of screens, and test whether it still reads as on-brand without a fixed layout.
What a designer's job becomes when you design the rules, not the screens
This is David's core reframe, and it's the one I'd bet sticks with people longest. When interfaces are generative, a designer can no longer specify every artifact. What you specify instead is constraints, priorities, permissible transformations, and ethical boundaries. His analogy is architecture: an architect doesn't dictate where every person stands in a building. They create a structure inside which a huge range of unpredictable behavior happens safely.
The old question was "where should this button go." The new one is "under what conditions should this button even exist." David calls that a profoundly different profession, and in his view, a harder design problem than the one it's replacing.
Playbook move: for your next feature spec, write the conditions under which it exists first. Wireframes come after.
Simulating users before you build (and why real interviews matter more, not less)
I asked David directly: can AI personas replace user interviews. No, and his answer is worth sitting with. "We simulate aircraft before we build them. That doesn't mean we stop using wind tunnels." Wind tunnels don't eliminate flight testing, because each method catches a different class of error. Reality gets the privilege of disagreeing with your model, and that disagreement is where the real learning happens.
What worries him about AI simulation specifically is that it's more convincing than the methods it's replacing, which makes it easier to mistake plausibility for evidence. A generated persona will produce a detailed, confident explanation for why "Jane, 34" abandoned checkout. Jane doesn't exist. That's not research, it's a well-written guess. Used correctly, simulation generates hypotheses and predicts behavioral distributions. Real interviews then become more valuable, not less, because you're hunting for exactly where reality diverges from the model. That divergence is usually where good discovery starts.
Playbook move: run the simulated pass first to generate hypotheses, but log every place a live interview or usability test contradicts it. That gap list is your actual research backlog.
Getting past "which version won" to actually explaining why
David walked through a concrete example instead of an abstract one: modeling decision fatigue with QFIT and finding it follows an interference pattern, the same math behind noise-canceling headphones, where competing options create literal wave-cancellation that leaves a user stuck in non-action. UX teams normally just label that anecdotally: "too many options, people struggle to choose." Having an underlying model explains it instead of describing it, and predicts when it'll happen again. He's clear this is one validated example, not proof the whole framework is correct, and he means it when he says he wants people to try to break it.
Playbook move: when a test result surprises you, don't stop at "variant B won." Write down the mechanism you think explains it, then check whether that mechanism predicts the result of a different test you haven't run yet.
The line between nudging and manipulating
I pointed out that David's healthcare context gives him cleaner ethics than most of us get. His answer generalized further than I expected. Persuasion itself isn't immoral, parents, doctors, and teachers all persuade. Manipulation starts when you understand someone's cognitive state and use that understanding to diminish their agency instead of restore it. In his words: if I know you're confused and use that to help you make an informed decision, I've increased your agency. If I know you're confused and exploit it to make the decision I want easier than the one you'd otherwise choose, I've diminished it.
AI raises the stakes on this because a system can understand the conditions driving your behavior better than you consciously understand them yourself. That's not a hypothetical. It's already true.
Playbook move: for any AI-personalized flow, run one test: does this use what we know about the user's state to help them decide, or to make one option artificially easier? Kill or redesign anything that fails it.
One thing to try this week
Neither of us handed the audience a single homework assignment, but the throughline of the whole conversation compresses into one sprint-sized habit: add a mechanism hypothesis to how your team already writes tickets, before you build anything AI-adaptive or not.
Playbook move: add two lines to your next sprint ticket template. "Mechanism hypothesis: ___" and "This would be falsified if: ___." It's the lightest version of the shift from gut feel to science, and it costs you nothing to start.
Where Evermuse fits
Everything in this conversation gets harder without a live picture of what users actually do and say. Evermuse exists to help teams scale their listening and their empathy as they build faster with AI, because staying connected to real users only gets harder as build speed goes up. David said something on the call I didn't prompt: he personally uses Evermuse as the bridge across his own stack, ChatGPT, Cursor, Figma, FigJam, Miro. That's the use case in one sentence. Not a replacement for interviews and usability testing. The thing that makes sure your model of the user keeps getting checked against reality instead of drifting away from it.
If you want to try it, install the MCP at evermuse.com/mcp.
About the speakers
David Sternberg is Design Principal at NiaHealth and the author of The Flow Equation: Quantum-Fluid Mechanics of User Experience. He spent five years at HelloFresh, most recently as Global Director of Product Design across the US, Canada, and Europe, and before that led UX and product design at OSRAM and creative direction at GE.
Eran Dror is the Co-Founder & CEO of Evermuse, where he ships complex production code with AI every day. He is the Managing Partner of Remake Ventures, has helped 40+ startups raise $300M+ by finding product-market fit, and had his first exit with SetJam (acquired by Motorola in 2012). Connect with him on LinkedIn.