Skip to content
View ariscaga's full-sized avatar
🤔
Building something cool
🤔
Building something cool

Block or report ariscaga

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ariscaga/README.md

Human-AI Evaluation Banner

Hi, I’m Ariana (@ariscaga)

Founder of QuorLoop | Building human–AI evaluation infrastructure

Before people interact with your AI, include them first. I build systems that help AI teams evaluate whether people can use, understand, and appropriately rely on their products—not just whether the model produces a good answer.

An AI response can be technically correct while leaving someone unsure what it means or what to do next. My work focuses on those interactions: confusion, task abandonment, misplaced reliance, and difficulty recognizing when human help is needed. Through QuorLoop, I’m building infrastructure that connects scenario testing, agentic simulation, and community evaluation to product decisions—and helps teams test whether their changes improve the experience.


Current Focus

QuorLoop — Evaluation for Consequential AI

QuorLoop helps teams evaluate specific human–AI interactions in products that shape access to healthcare, public services and benefits, legal services, adult education, and work. Our approach combines:

  • Realistic scenarios: Situations grounded in users’ goals, constraints, missing information, and decisions.
  • Agentic simulation: Predicted user responses that help explore possible interaction problems and prioritize questions for testing.
  • Relevant human evaluation: Structured testing with people whose experiences fit the selected workflow.
  • Comparative evaluation: Assessing whether a proposed change improves the interaction. The goal is to turn findings into concrete product changes and evidence about whether those changes help.

Community AI Evaluator Network

We’re building a paid community AI evaluator network in New York City, training people to contribute their perspectives through structured, mobile-first evaluation tasks.

Contributors help us examine how people interpret AI responses, choose their next steps, and recognize when they need more information or human support.

Their participation brings community experience into AI development and creates paid opportunities to contribute to how these products are evaluated.


Human × Agentic Evaluation

I’m developing methods to compare predicted behavior from agentic simulations with observed responses in human evaluations.

We call this difference the Signal Gap: where simulated expectations diverge from what participants understand or do in a defined test. My research interests include cohort behavior estimation, simulation calibration, and whether findings transfer across workflows. Simulation predictions require human comparison; results from one cohort or test do not automatically generalize to another.


How I Build

  1. Start with the interaction
  2. Define who the product serves, what they are trying to accomplish, and the decisions, constraints, and failure conditions involved.
  3. Include people early
  4. Bring relevant people into evaluation before launch and continue learning as the product changes.
  5. Evaluate beyond technical correctness
  6. Examine whether people understand a response, can take an appropriate next step, and know when to verify it or seek human help.
  7. Connect simulated and observed behavior
  8. Use simulations to explore hypotheses. Use human evaluation to observe responses in the test context and examine where predictions hold or break down.
  9. Test the change
  10. Translate findings into proposed improvements, then evaluate whether the revised interaction performs better.
  11. Make evidence reusable
  12. Preserve scenarios, methods, findings, and their limits so each evaluation can inform future work without overstating what transfers.

Background

  • Second-time founder; founder of QuorLoop
  • Previously founded DivySci, a multimillion-dollar company serving more than 10,000 users
  • Blue Ridge Labs fellow at the Robin Hood Foundation
  • M.S. in Information and Knowledge Strategy, Columbia University
  • B.S. in Computer Science, Pace University
  • Two-time NSF-funded founder
  • Support across my founder journey from Google for Startups, AWS, Camelback Ventures, Black Ambition, and the Roddenberry Foundation

Technical Context

AI and Evaluation

LLMs · RAG · Agentic Systems · Scenario Testing · Human Evaluation · Simulation Calibration · Comparative Evaluation

Backend

Python · Node.js · PostgreSQL · Supabase

Frontend

React · TypeScript · Next.js

Product and Systems

Human–AI Interaction · AI Evaluation Infrastructure · Cohort Behavior Estimation · AI Product Strategy · Product Architecture


Core Themes

Human-Centered AI Evaluation · Human × Agentic Evaluation · AI Safety · Trust and Decision Risk · Human Signal Infrastructure · Responsible Model Behavior


Philosophy

The people an AI product is meant to serve should help shape how it is evaluated.

I build the systems that make that participation useful for product decisions.

Connect

Website: www.arianaabramson.com

LinkedIn: http://linkedin.com/in/arianaabramson

Email: ariana.abramson@gmail.com

Pinned Loading

  1. d15-docs d15-docs Public

    Strategic doctrine, product thesis, and ecosystem architecture for D15 — human intelligence infrastructure for the AI era.

    1

  2. d15-contributor d15-contributor Public

    Mobile-first contributor platform for learning, earning, and participating in the AI economy.

    1

  3. d15-b2b d15-b2b Public

    Enterprise human intelligence infrastructure for AI systems — task generation, governed review pipelines, model integration, and continuous evaluation.

    1

  4. ai-infrastructure-framework ai-infrastructure-framework Public

    Technical showcase of the SCALE FACTOR™ Assessment methodology - a proprietary scoring framework for measuring AI infrastructure readiness across 7 critical dimensions. Includes scoring logic, samp…

    TypeScript 1

  5. ai-social-assessment ai-social-assessment Public

    Evaluate AI initiatives for social value, ethics & community impact. 500-point scoring system across 5 dimensions. Try it: social-aivalue.netlify.app