Exploring ride-sharing usage and attitudes with conversational surveys
A product lead at a ride-hailing company is holding a stack of AI-moderated interview transcripts, each one long, candid, and full of opinions about fares, wait times, and safety. Riders will talk at length in a chat window; that part is settled. What the transcripts cannot settle is which of those things, if changed, would actually move a ride request. Answering that requires a randomized experiment, not another round of open-ended conversation.
- Conversational AI interviews produce longer, more candid answers about ride-sharing price, wait time, and safety than click-through surveys, but word count measures engagement, not causal effect.
- Only a randomized discrete-choice experiment, estimated with a method like Mixed Logit or ICLV, isolates which specific attribute (price, wait time, or a named safety signal) changes an actual ride request.
- Stated-preference research carries a documented say-do gap: hypothetical bias systematically inflates stated intent and willingness-to-pay relative to real transactions (ScienceDirect).
- Adoption context matters when reading these studies: 36 percent of U.S. adults had used ride-hailing by fall 2018, up from 15 percent in 2015, but only about one in ten users ride weekly (Pew Research Center), the newest national figures Pew has published.
- Simulated experiments can be checked against real human studies before a buyer trusts one: Subconscious reports 93 percent replication accuracy (go.subconscious.ai/paper), a validation-set result, not a guarantee for an untested market.
More talk is not more proof
AI-moderated conversational interviews are displacing static click-through surveys for ride-sharing attitude research. Outset runs AI-moderated interviews across video, voice, and text (outset.ai), one example of the shift toward respondents typing or talking instead of clicking through a five-point scale. Ride-sharing researchers have followed the category: a rider typing through what safety and convenience mean to them produces far richer raw material than a Likert scale ever did. That richness is real. What it proves is a separate question.
Do longer answers mean better ride-sharing insights?
No. A longer transcript is evidence that someone engaged with the question, not evidence that their answer predicts what they would actually do. Stated-preference research has spent decades documenting hypothetical bias: meta-analytic work shows it is pervasive across discrete-choice and willingness-to-pay studies, and it systematically inflates stated intent relative to real transactions (ScienceDirect). Conversational formats can make this worse, not better. A free-form chat invites social-desirability signaling: a rider typing about safety concerns is partly answering the question and partly performing a version of themselves they want the interviewer, human or AI, to see. A transcript that says "I care a lot about driver background checks" does not tell a product team whether adding a visible background-check badge would change one ride request. That gap between what people say and what they do is exactly what a randomized experiment is built to close.
What the ICLV rideshare study shows, and where it stops
A 2025 U.S. national study applied an Integrated Choice and Latent Variable (ICLV) model to 8,296 survey responses to map latent attitudes (safety perception, service experience, time sensitivity, and environmental awareness) onto interest in pooled rideshare (MDPI). ICLV is a genuine step up from a plain attitude survey: it models the latent construct (how much someone actually weighs safety) rather than just tallying who mentioned safety. But ICLV is an estimator, not a causal method on its own. It tells you how latent attitudes correlate with stated interest in pooled rides across a sample. It does not, by itself, tell a buyer that raising the visibility of a specific safety feature, holding price and wait time constant, would change real ride requests. That identification only comes from randomizing the attribute itself inside the experiment design, not from modeling attitudes more precisely after the fact.
Which ride-sharing attribute actually changes a ride request?
Only a randomized manipulation of the attribute itself answers that, run through a discrete choice model. The design: show respondents repeated trade-offs, ride A at one price, wait time, and safety-signal combination versus ride B at another, with the levels assigned at random, then estimate the results with McFadden discrete choice, Mixed Logit, or ICLV. These are estimators; the causal claim comes from the randomization in the design, not from the estimator's name. Mixed Logit is worth the extra complexity here because a plain multinomial logit assumes independence of irrelevant alternatives (IIA), meaning it assumes a new safety feature pulls share from price-sensitive and safety-sensitive riders in fixed proportion. Ride-hailing riders don't split that cleanly, so a flat logit model can misstate which attribute substitutes for which.
Why does adoption data matter to this decision?
Because most people a ride-sharing survey reaches are occasional users, and their stated attitudes are more likely to diverge from behavior than a daily rider's would. Pew found 36 percent of U.S. adults had used a ride-hailing service by fall 2018, up from 15 percent in late 2015, but only about one in ten users ride weekly (Pew Research Center). Those are the newest national figures Pew has published; the specific percentages are now eight years old, and Pew has not published a newer national breakdown, so whether the trial-heavy, weekly-light pattern still holds is unverified. A conversational interview about safety or price run against this population is mostly talking to people whose "I would ride more if..." statements have never been tested against a real fare or a real wait. That's not a reason to skip attitude research; it's a reason to treat it as a hypothesis generator and send the resulting hypotheses into a randomized test before a pricing or safety-feature decision gets made on the strength of a transcript.
Conversational interview or randomized experiment: how to pick
| Conversational AI interview | Randomized discrete-choice experiment | |
|---|---|---|
| What it captures | Open-ended language about price, wait time, and safety; which topics riders bring up unprompted | A forced trade-off between specific attribute levels, e.g., a fare change vs. a wait-time change vs. a named safety feature |
| What it can't tell you | Whether a stated concern would change an actual ride request | The respondent's own reasoning, unless paired with a follow-up open question |
| Underlying method | Text, voice, or video chat, coded by theme | McFadden discrete choice, Mixed Logit, or ICLV, estimated on randomized choice tasks |
| Best for | Surfacing candidate attributes and language before a test is designed, not sizing an intervention | Deciding whether to ship a price, wait-time, or safety change, not describing rider sentiment |
How do you validate a causal estimate before shipping it?
Check it against a real human study before trusting it on a new market. Subconscious reports 93 percent replication accuracy on its validation set (how often a simulated study reproduces the direction and outcome of the original human study) (go.subconscious.ai/paper). That number describes past replications, not a guarantee for a market that hasn't been tested. Published studies can sit inside a model's training data, so replication against a known study is a weaker check than replication against a genuinely new holdout. The protocol behind that figure is built to address that gap, though the number alone does not. The leaderboard publishes replication results across methods and categories so a buyer can see how a given estimator performs before committing budget to a live test. Any confidence interval that comes out of a simulated discrete-choice experiment covers the estimated effect within that simulated population; it does not, by itself, bound the real market without a validated replication behind it.
The decision this article owns
Don't let a conversational interview stand in for a pricing, wait-time, or safety-feature decision. Use it to generate candidate attributes and language, then run a randomized discrete-choice experiment on those specific attributes before shipping. Start by pulling the two or three attributes your last round of AI interviews mentioned most (price, wait time, a named safety signal) and write them up as levels for a choice task; the case studies page has worked examples of that step. If you want a second opinion on the design before you run it, meet with the team.