- Question
- Are chatbots useful?
- Position‹3 of 3
- It depends — useful for bounded tasks with verification
- Argument1 of 2›
The useful zone is a 'jagged frontier' the user must police
The argument
Mollick and colleagues' 2023 BCG study described chatbot capability as a 'jagged frontier' — a landscape of tasks where the model performs strongly on some and unexpectedly badly on others, with the boundary visible only to those who already know the right answer. Inside the frontier, consultants using GPT-4 produced higher-quality work measurably faster. Outside the frontier, on a deliberately chosen task designed to expose the model's blind spots, consultants who trusted the chatbot performed worse than those without it; they were confidently led to a wrong answer that those working unaided either avoided or caught. The implication is awkward. To benefit reliably from a chatbot, a user must have enough domain expertise to know when the model has wandered off the frontier — but if the user has that expertise, they need the chatbot less. The most enthusiastic adopters tend to be those least equipped to identify its mistakes, while the people who could safely identify them tend to be those with the least incentive to delegate. Chatbots are therefore useful, but the conditions under which their utility is reliable are narrower than the marketing or the headline studies suggest, and the unattended use case — confident non-experts producing decision-grade output — is the one in which they cause the most damage.
Premises
Counter-arguments
Enthusiasts reply that the 'jagged frontier' is real but shrinking, and that the expertise paradox is overstated. Users need not be domain experts to catch errors when the workflow supplies external verification — running code, checking citations, cross-referencing a source — so utility does not depend on already knowing the answer. Newer models have receded many of the blind spots the 2023 study exposed, and the same 'confidently wrong' risk applies to search engines, textbooks and human colleagues, which we nonetheless use productively with sensible checking. The BCG study also measured a single task deliberately chosen to defeat the model, whereas typical use is not adversarially selected. Chatbots, on this view, are simply useful with ordinary verification, not narrowly useful only to those who least need them.
Rejecting the premises
[Rejecting P1] External verification — running code, checking sources, cross-referencing — lets non-experts catch errors without already knowing the answer, so reliable utility does not require the domain expertise the argument says is needed. [Rejecting P2] The BCG task was deliberately chosen to expose blind spots and predates rapid model improvement, so the 'confidently wrong outside the frontier' finding overstates the risk in typical, non-adversarial use.