AI advice in the real world

When AI gives advice, correct is not enough.

People increasingly ask large language models what to do about security, privacy, fraud, stalking, harassment, domestic violence, and other high-stakes situations. The answer may sound confident. That does not make it safe, well-sourced, or useful.

A good answer must come from the right sources, fit the user’s actual situation, and lead to actions that are safe and feasible.
Read our papers
Stylized illustration of a worried person asking an AI chatbot for help with stalking and harassment, while the chatbot gives exaggerated, overconfident VPN advice.
A deliberately exaggerated example: generic security advice can sound reassuring while missing the actual threat, the user’s context, and the consequences of acting on it.
The answer can be wrong.

Models may reproduce stale guidance, incomplete claims, or product marketing instead of authoritative evidence.

The answer can be right, but unsafe.

Technical correctness can fail when the model ignores who controls the device, who may be adversarial, or what happens next.

The answer can be unreadable or unusable.

High-stakes advice must be prioritized, clear, actionable, and appropriate for the person receiving it.

Three problems AI builders need to solve

Correctness, context, and actionability are tightly connected.

These are not separate product niceties. In high-stakes advice, failure in any one of them can make the overall response misleading or unsafe.

1

Context-aware safety

A technically valid recommendation can be the wrong recommendation for this person, in this situation.

For example, in technology-facilitated abuse or domestic violence situations:

  • Who controls the device, account, network, or data?
  • Is access consensual, shared, compromised, or adversarial?
  • Could a change expose intent, trigger retaliation, destroy evidence, or close off safer options?
2

Authoritative correctness

Finding “an answer” is not enough. Models need to identify which evidence deserves trust for the specific question.

  • Prefer first-party documentation, government guidance, and relevant academic research.
  • Distinguish evidence from SEO content, forums, outdated guidance, and product marketing.
  • Use the source that matches the actual product, threat model, version, and context.
3

Actionable delivery

Advice must be possible to follow, not merely readable or technically sophisticated.

  • Prioritize the safest next steps instead of producing a long generic checklist.
  • Give accurate step-by-step instructions when procedural detail matters.
  • Adapt tone, detail, and urgency to the user’s constraints and situation.
The central evaluation question should be: “What happens if the user actually follows this advice?”
That requires source quality, situational reasoning, and realistic actionability, not factuality alone.
From answer quality → outcome quality
Why this generalizes

The problem is bigger than any one safety domain.

Technology-facilitated abuse makes the failure mode unusually visible because the environment may contain an active adversary. But the same design problem appears whenever advice is consequential, the evidence landscape is noisy, or the user’s circumstances change what “good advice” means.

Fraud & scamsStalking & harassmentElder financial exploitationAccount compromisePersonal security & privacyCybersecurity incident responseOther high-stakes advice
AI systems should not only ask, “Is this statement true?” They should also ask, “Is it true here, for this user, from a trustworthy source, and is acting on it safe?”
Research behind the argument

Two studies, one broader lesson.

Across end-user security and technology-facilitated abuse, we repeatedly found that plausible LLM responses can fail on accuracy, sourcing, context, safety, and actionability.

ACSAC 2025

Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions

Vijay Prakash, Kevin Lee, Arkaprabha Bhattacharya, Danny Yuxing Huang, Jessica Staddon

An expert evaluation of responses to 900 end-user security questions across seven security areas. The study identifies recurring problems including stale and incomplete guidance, over-prescription of security products, inaccurate procedural steps, and responses that appear to overstate technology capabilities rather than reflect current authoritative guidance.

USENIX Security 2026
Runner-up · Distinguished Paper Award

Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse

Vijay Prakash, Majed Almansoori, Donghan Hu, Rahul Chatterjee, Danny Yuxing Huang

An expert-led evaluation of LLM responses to real-world technology-facilitated abuse questions, combined with feedback from people with lived experience. The study extends response quality beyond accuracy and completeness to safety and actionability, showing why context, power dynamics, user constraints, and downstream consequences matter.

People

The research team and collaborators.

The work brings together researchers across academia and industry.

Core researchers

Vijay Prakash, PhD

Graduated with a PhD from NYU in 2026 and is now a postdoctoral fellow at Northeastern University.

Danny Yuxing Huang, PhD

Assistant Professor at New York University and principal investigator.

Collaborators
Majed Almansoori
University of Wisconsin–Madison
Arkaprabha Bhattacharya
Cornell Tech
Rahul Chatterjee, PhD
Associate Professor, University of Wisconsin–Madison
Donghan Hu, PhD
Postdoctoral Fellow, New York University
Kevin Lee, PhD
JPMorgan Chase
Jessica Staddon, PhD
Professor, Northeastern University
Work with us

Interested in improving AI advice?

We would be glad to talk with AI model, safety, product, security, and nonprofit teams about evaluation, datasets, context-aware safeguards, authoritative retrieval, and other ways to make high-stakes AI advice safer and more useful.

Contact the Research Team