Sid Mehta Works
Obin

Obin

Project details

Year

2026

Client

Obin · concept

Role

Design · Prototype

Tools

Claude · HTML · CSS · JS

Live

Run the prototype

Project highlights

38

figures in one memo

28

proved without a person

2

drawn back at random, unnamed

Overview

A verification workspace for AI-written venture debt memos. The agent drafts; the interface makes visible which figures it proved against a source, which it could not, and which two of its own it drew back at random for a person to check.

Agentic AI · FintechDesign · Prototype · Obin · concept

Obin's agent drafts a venture debt credit memo end to end. The draft is fluent, and fluency is the problem: a model that writes well makes a wrong number look exactly like a right one. This tool cannot stop a bad number. It can stop one from looking like a good one.

The whole run, sped up. It reads the data room, writes the memo, marks every figure it asserts as it lands, and then stops and names the four things it will not decide.

The speed is not the achievement. Four minutes to a drafted memo only matters if a person can tell, without rereading the data room, which of its numbers were proved and which were asserted well. Everything below is that distinction, made visible.

Open the Obin prototype

The verification workspace, running in this browser. Open any figure to see what it was proved against, or what it was not.

Serif for what a person wrote. Monospace for what the machine asserted. The underline carries the state, so it survives being printed in black and white.

A credit memo where every machine-asserted figure is set in monospace with a coloured underline: blue for proved, amber for conflict, red for unsupported

Ordering is the product. Sorted by leverage crossed with reliability, the undisclosed lien is first. In the order a person reading top to bottom would meet it, it is ninth.

The queue in risk order, with the senior blanket lien at the top marked unsupported
The same queue in document order, where the lien falls to ninth

Every claim opens into what the record says, what the agent did with it, and where the agent went when the data room came up short.

The verification pane: registry filing on the left, what the agent did with it on the right, and a note that the company did not supply the document
The header strip: time on the memo against a firm baseline, 38 figures, 28 proved without you, 12 to review

Model confidence appears nowhere in the product. A score is 99 when a tool ran the arithmetic and the model never touched the number, and 77 when it read the figure off a slide someone wrote to persuade you. Both are backtested against two hundred and fourteen closed memos where the analyst's own number exists to compare against. Probability measures fluency, and a hallucinated figure often scores higher than a correct hedged one.

The walkthrough page, stating the thesis and listing the four assumptions the design rests on
Reflection

The thing I keep coming back to is that the agent made the analyst’s job worse before it made it better. Four hours became four minutes, and what came out the other side was a well written memo with thirty eight numbers in it and no way to tell which the machine had invented. So the analyst checked all thirty eight. The speed went somewhere, but not to them. What I would test first is the one assumption the whole ordering rests on: that leverage crossed with reliability is the right sort. Every workflow decision here is inferred rather than observed, and the design says so on its own face rather than hiding it in a footnote.

Back to top
Next project Ledger A reviewer has twenty three submissions, one sitting, and a dozen browser tabs per cand... Back to top