Obin
Project details
Project highlights
38
figures in one memo
28
proved without a person
2
drawn back at random, unnamed
A verification workspace for AI-written venture debt memos. The agent drafts; the interface makes visible which figures it proved against a source, which it could not, and which two of its own it drew back at random for a person to check.
Agentic AI · FintechDesign · Prototype · Obin · concept
Obin's agent drafts a venture debt credit memo end to end. The draft is fluent, and fluency is the problem: a model that writes well makes a wrong number look exactly like a right one. This tool cannot stop a bad number. It can stop one from looking like a good one.
The whole run, sped up. It reads the data room, writes the memo, marks every figure it asserts as it lands, and then stops and names the four things it will not decide.
The speed is not the achievement. Four minutes to a drafted memo only matters if a person can tell, without rereading the data room, which of its numbers were proved and which were asserted well. Everything below is that distinction, made visible.
The verification workspace, running in this browser. Open any figure to see what it was proved against, or what it was not.
Serif for what a person wrote. Monospace for what the machine asserted. The underline carries the state, so it survives being printed in black and white.
Ordering is the product. Sorted by leverage crossed with reliability, the undisclosed lien is first. In the order a person reading top to bottom would meet it, it is ninth.
Every claim opens into what the record says, what the agent did with it, and where the agent went when the data room came up short.
Model confidence appears nowhere in the product. A score is 99 when a tool ran the arithmetic and the model never touched the number, and 77 when it read the figure off a slide someone wrote to persuade you. Both are backtested against two hundred and fourteen closed memos where the analyst's own number exists to compare against. Probability measures fluency, and a hallucinated figure often scores higher than a correct hedged one.
Back to top↑The thing I keep coming back to is that the agent made the analyst’s job worse before it made it better. Four hours became four minutes, and what came out the other side was a well written memo with thirty eight numbers in it and no way to tell which the machine had invented. So the analyst checked all thirty eight. The speed went somewhere, but not to them. What I would test first is the one assumption the whole ordering rests on: that leverage crossed with reliability is the right sort. Every workflow decision here is inferred rather than observed, and the design says so on its own face rather than hiding it in a footnote.
Next project Ledger→ A reviewer has twenty three submissions, one sitting, and a dozen browser tabs per cand... Back to top↑