A small team with one belief.
Self-report is uncorrelated with reality. The only way to know what an AI is doing to people is to watch them, not ask them.
The gap between two answers.
BIE was started in 2026 to fill the gap between “the model is fine” and “the user is fine.” Those are two different questions, and the industry had only built an instrument for the first.
Eval tools answered whether the output was faithful, safe, fast, and cheap. That work matters. But it stops at the model boundary. Nobody was measuring what the output did to the person reading it, turn by turn.
Eval tools answered the first. We built the instrument for the second.
So we built the second instrument. It watches the human side of the conversation on dimensions nobody is watching for yet: trust calibration, frustration buildup, dependency drift, silent abandonment. The deployment is the subject. Individuals are evidence.
No customer logos we can't stand behind. No fabricated findings. No clinical language about the people we measure. The research is the evidence. Read more about how we operate on the about page.
Want to talk?
Questions about the engine, the research, or where this is going. One call, no pitch deck.