Before people interact with your AI, include them first. I build systems that help AI teams evaluate whether people can use, understand, and appropriately rely on their products—not just whether the model produces a good answer.
An AI response can be technically correct while leaving someone unsure what it means or what to do next. My work focuses on those interactions: confusion, task abandonment, misplaced reliance, and difficulty recognizing when human help is needed. Through QuorLoop, I’m building infrastructure that connects scenario testing, agentic simulation, and community evaluation to product decisions—and helps teams test whether their changes improve the experience.
QuorLoop helps teams evaluate specific human–AI interactions in products that shape access to healthcare, public services and benefits, legal services, adult education, and work. Our approach combines:
- Realistic scenarios: Situations grounded in users’ goals, constraints, missing information, and decisions.
- Agentic simulation: Predicted user responses that help explore possible interaction problems and prioritize questions for testing.
- Relevant human evaluation: Structured testing with people whose experiences fit the selected workflow.
- Comparative evaluation: Assessing whether a proposed change improves the interaction. The goal is to turn findings into concrete product changes and evidence about whether those changes help.
We’re building a paid community AI evaluator network in New York City, training people to contribute their perspectives through structured, mobile-first evaluation tasks.
Contributors help us examine how people interpret AI responses, choose their next steps, and recognize when they need more information or human support.
Their participation brings community experience into AI development and creates paid opportunities to contribute to how these products are evaluated.
I’m developing methods to compare predicted behavior from agentic simulations with observed responses in human evaluations.
We call this difference the Signal Gap: where simulated expectations diverge from what participants understand or do in a defined test. My research interests include cohort behavior estimation, simulation calibration, and whether findings transfer across workflows. Simulation predictions require human comparison; results from one cohort or test do not automatically generalize to another.
- Start with the interaction
- Define who the product serves, what they are trying to accomplish, and the decisions, constraints, and failure conditions involved.
- Include people early
- Bring relevant people into evaluation before launch and continue learning as the product changes.
- Evaluate beyond technical correctness
- Examine whether people understand a response, can take an appropriate next step, and know when to verify it or seek human help.
- Connect simulated and observed behavior
- Use simulations to explore hypotheses. Use human evaluation to observe responses in the test context and examine where predictions hold or break down.
- Test the change
- Translate findings into proposed improvements, then evaluate whether the revised interaction performs better.
- Make evidence reusable
- Preserve scenarios, methods, findings, and their limits so each evaluation can inform future work without overstating what transfers.
- Second-time founder; founder of QuorLoop
- Previously founded DivySci, a multimillion-dollar company serving more than 10,000 users
- Blue Ridge Labs fellow at the Robin Hood Foundation
- M.S. in Information and Knowledge Strategy, Columbia University
- B.S. in Computer Science, Pace University
- Two-time NSF-funded founder
- Support across my founder journey from Google for Startups, AWS, Camelback Ventures, Black Ambition, and the Roddenberry Foundation
AI and Evaluation
LLMs · RAG · Agentic Systems · Scenario Testing · Human Evaluation · Simulation Calibration · Comparative Evaluation
Backend
Python · Node.js · PostgreSQL · Supabase
Frontend
React · TypeScript · Next.js
Product and Systems
Human–AI Interaction · AI Evaluation Infrastructure · Cohort Behavior Estimation · AI Product Strategy · Product Architecture
Human-Centered AI Evaluation · Human × Agentic Evaluation · AI Safety · Trust and Decision Risk · Human Signal Infrastructure · Responsible Model Behavior
I build the systems that make that participation useful for product decisions.
Website: www.arianaabramson.com
LinkedIn: http://linkedin.com/in/arianaabramson
Email: ariana.abramson@gmail.com


