Conversational Finance using Semantic AI
Let's be honest: finance does not have a data problem. Finance has a trust problem.
Every team I know that builds finance software carries the same scar. A customer asks the product for revenue, the product answers, and the number in the board deck is not the number their accountant signed. Nobody lied. The customer and the product simply meant two different things by the same word.
Now put an AI in the middle of that conversation. It is fluent, it is fast, it is confident, and it has never once read your customer's revenue recognition policy.
ScramDB fixes this where the numbers live. Every ScramDB comes with Semantic AI built in: install the database and it is already running, ready for Claude or for the agent inside your own product.
Plausible is not correct
Here is a question every finance tool claims to answer: what was our revenue this quarter?
I loaded a small ledger into ScramDB and asked it two ways. The first way is what a language model writes when you hand it raw tables and a prompt. It finds a column called account_type, filters on 'revenue', sums it, and moves on. The second way is the governed metric, the one finance actually signed off on, which is net of refunds.
What a model writes against raw tables
gross revenue: account_type = 'revenue'
SELECT sum(signed_amount) FROM gl_lines WHERE account_type = 'revenue' AND posted >= '2026-07-01'The governed metric (net_revenue)
net of refunds, as finance defined it
SELECT sum(CASE WHEN (account_type) IN ($1, $2) THEN (signed_amount) END) AS net_revenue ...€30,000 of refunds counted as revenue
| Query | Result |
|---|---|
| What a model writes against raw tables | €724,000 |
| The governed metric (net_revenue) | €694,000 |
Both queries run. Both return a clean, plausible number. Neither throws an error. One of them is thirty thousand euros wrong, and you will not find out from the SQL, because the SQL is fine. The definition is what went missing.
This is the failure mode nobody puts on the slide. A text-to-SQL agent does not fail loudly. It fails politely, with a number that looks exactly like the right one. It picks the wrong join, it invents what "revenue" means, it quietly counts a refund as income, and the answer sails straight into your customer's board pack. With your product's name on it.
The fix is not a smarter model. The fix is to stop letting the model decide what your words mean.
Semantic AI, built into the database
ScramDB carries the meaning of your numbers itself. Semantic AI comes with every ScramDB: not an add-on, not a separate service, not a second vendor. It gives an agent the three things that separate a trustworthy answer from a plausible one:
- A semantic layer. What your metrics and dimensions mean, and exactly how they are calculated.
net_revenueis revenue plus refunds, and refunds are negative.gross_profitis net revenue plus cost of revenue.entityis a dimension you can slice both by. Written down once, in ordinarysemantic.*tables inside ScramDB, replicated with the rest of your data and governed byGRANTlike any other table. - Hints. Usage guidance on every metric and dimension: not only what it means, but when to reach for it and when not to. "Revenue always means net unless the user explicitly says gross." "Prefer quarter grain for trends." The agent reads them before it answers.
- Instructions. The rules for how the agent behaves with your data: which assumptions are reasonable, what to do with an ambiguous question, and when to stop and ask. "Fiscal year is the calendar year." "If a question could mean two metrics, ask which one." "If a metric is missing, say so and propose its definition, never approximate it with raw SQL." They arrive in the MCP handshake, before the agent's first tool call.
A semantic layer tells the agent what a number is. Hints tell it when to use it. Instructions tell it what to do when the question is not clear. Skip the instructions and the agent guesses whenever a question is vague. With them, it asks.
And it remembers. When someone on the finance team tells the agent that revenue means net of refunds, or that a vague question deserves a clarifying one, one tool call saves it: define_model or define_instructions. ScramDB keeps it in its own semantic.* tables, and from then on every conversation, for every user and every agent, starts from what the team already taught it. Who is allowed to teach it is a GRANT, like everything else.
And it refuses to rot. Every definition is validated against the live catalog the moment it is written. Rename a column your metric depends on and ScramDB tells you, loudly, instead of letting a spreadsheet of YAML drift out of sync with reality for six months.
ScramDB
Three properties make this safe enough for money:
- Filters are data, not SQL. The agent sends
{"field": "entity", "op": "eq", "value": "EU"}. Values travel as bound parameters. There is no string for anyone to inject into, by construction. - Every query runs as the caller. Turn on
MCP_AUTH=basicand each request executes as that person's own role in ScramDB. Row-level security applies. The EU controller asking the same question as the group CFO gets exactly the rows her policy allows, and not one more. In a multi-tenant product, that is tenant isolation you did not have to write. - Every answer carries its receipt. Rows, the exact SQL, the parameters, and an honest
truncatedflag. "Where did that number come from?" has a one-line answer.
And because this is ScramDB, the semantic server itself runs inside the database, in the same sealed Afterburner sandbox that runs every package: deny by default, no filesystem, no ambient network, nothing it was not explicitly granted.
Governed answers in conversation
Enough architecture. Here is the conversation your product can hand to its CFO users. This one is Claude, connected to ScramDB over MCP, against a three-entity ledger (EU, NA, AP) with two quarters of postings.
describe arguments
{}2 models · 6 metrics · 5 instructions. Hint on net_revenue: THE revenue metric, net of refunds
query arguments
{
"metrics": [
"net_revenue"
],
"dimensions": [
"ledger.entity"
],
"filters": [
{
"field": "posted",
"op": "between",
"value": [
"2026-07-01",
"2026-09-30"
]
}
]
}| entity | net_revenue |
|---|---|
| AP | €146,000 |
| EU | €234,000 |
| NA | €314,000 |
Exact SQL
SELECT (entity) AS entity, sum(CASE WHEN (account_type) IN ($1, $2) THEN (signed_amount) END) AS net_revenue FROM gl_lines WHERE (posted) >= $3 AND (posted) <= $4 GROUP BY entity LIMIT 101
params: ["revenue","refund","2026-07-01","2026-09-30"]
batch arguments
{
"requests": [
{
"tool": "query",
"arguments": {
"metrics": [
"opex"
],
"dimensions": [
"ledger.entity",
"vendor"
],
"filters": [
{
"field": "posted",
"op": "between",
"value": [
"2026-07-01",
"2026-09-30"
]
}
]
}
},
{
"tool": "query",
"arguments": {
"metrics": [
"budget"
],
"dimensions": [
"opex_budget.entity"
],
"filters": [
{
"field": "period",
"op": "between",
"value": [
"2026-07-01",
"2026-09-30"
]
}
]
}
}
]
}opex (non-null only)
| entity | vendor | opex |
|---|---|---|
| EU | Northwind Cloud | -€52,500 |
| NA | Contoso Ads | -€47,000 |
| AP | Contoso Ads | -€16,000 |
budget
| entity | budget |
|---|---|
| EU | -€45,000 |
| NA | -€50,000 |
| AP | -€20,000 |
From ScramDB's governed net_revenue metric, compiled to:
SELECT (entity) AS entity, sum(CASE WHEN (account_type) IN ($1, $2) THEN (signed_amount) END) AS net_revenue FROM gl_lines WHERE (posted) >= $3 AND (posted) <= $4 GROUP BY entity LIMIT 101
params: ["revenue","refund","2026-07-01","2026-09-30"]
Notice what the agent did not do
It did not
- Write SQL against gl_lines
- Guess what "revenue" means
- Invent a metric on the fly
It did
- 1describeRead that revenue means net of refunds
- 2queryAsked for net_revenue, a metric that already existed
- 3ScramDBCompiled the SQL from that definition
Then the follow-up every CFO actually asks: why did margin move?
| Entity | Q2 2026 | Q3 2026 |
|---|---|---|
| EU | 71.9% | 56.4% |
| NA | 67.6% | 41.7% |
| AP | 60.0% | 71.9% |
EU and NA both lost more than fifteen points of gross margin between Q2 and Q3, while AP went the other way. One query call, two governed metrics, grouped by entity and quarter. The story behind EU is sitting right there in the ledger: a 21,000 euro SLA credit on the revenue side, and hosting cost climbing on the other.
And the question that ruins a Friday afternoon: are we over budget anywhere?
| Entity | Actual | Budget | Variance |
|---|---|---|---|
| EU | €52,500 | €45,000 | €7,500 over |
| NA | €47,000 | €50,000 | €3,000 under |
| AP | €16,000 | €20,000 | €4,000 under |
The agent fetched actual opex and the budget in one round trip, using ScramDB Semantic AI's batch tool, and lined them up. EU is 7,500 euros over plan, and every euro of EU opex this quarter went to a single vendor. That is not a dashboard your team had to build last month. That is the ledger, as of the last committed posting, answering a question nobody anticipated.
Connecting an agent
This is the whole setup on the Claude Code side. The agent inside your own product connects the same way, and so does any MCP host that speaks streamable HTTP:
{ "mcpServers": { "scramdb-semantics": { "url": "http://127.0.0.1:9191/mcp" } } }
There is nothing to install and nothing to deploy. Semantic AI comes with ScramDB. It ships as a built-in package with autostart, so the database starts it at boot, supervises it, and serves MCP on port 9191. If you ran the ScramDB quick start, the -p 9191:9191 in that docker run was the entire installation.
Defining a metric is either a conversation ("define a model for the ledger, revenue is net of refunds") or one tool call:
{
"name": "net_revenue",
"agg": "sum",
"arg": "signed_amount",
"filter": { "field": "account_type", "op": "in", "value": ["revenue", "refund"] },
"hints": "THE revenue metric, net of refunds"
}
And when you are starting from nothing, suggest drafts a whole model from a table's schema in one call. Draft being the operative word: it happily offers you id_sum, because a machine can read a schema, but only your customer's finance team knows that refunds are negative and revenue means net. The schema is the machine's job. The meaning belongs to finance.
Matching payments to invoices
If you build finance software, you have watched e-invoicing go from optional to mandatory.
Country after country is making structured e-invoices mandatory, and more join every year. The invoice side of accounts payable is turning into clean, structured data, first inside each country and then across borders.
The bank side is not. Bank feeds still arrive shouting in capital letters, with a free-text payment reference somebody typed on their phone, and payment intermediaries sitting in the middle of half the transactions.
Matching the two is where the accounting hours actually go. Here is what that looks like when an AI does it on ScramDB, on six real-looking payments and eight open invoices from one client.
“NW-2026-0815 CUST 4711”
“1049384756 WGP.7361 PURCHASE AT CONTOSOADS”
“Invoice 2026 0921 less 2% early payment discount”
“Freight August”
“Account maintenance fee 08/2026”
“INV 2026/0930 + 2026/0935”
An amount-only matcher books this to Northwind. It is two Fabrikam invoices.
| Payment | Counterparty | Amount | Candidate invoice | Amount exact | Early-pay 2% | Supplier account | Reference in payment | Name in counterparty | Decision |
|---|---|---|---|---|---|---|---|---|---|
| 101 | NORTHWIND CLOUD | €1,190.00 | NW-2026-0815 | yes | no | yes | yes | yes | Match |
| 101 | NORTHWIND CLOUD | €1,190.00 | NW-2026-0901 | yes | no | yes | no | yes | Rejected: same amount and bank account, no reference |
| 102 | WOODGROVE PAY *CONTOSOADS | €2,380.00 | CA-77812 | yes | no | no | no | no | Needs a person: only the amount agrees |
| 103 | FABRIKAM OFFICE SUPPLY | €1,176.00 | 2026/0921 | no | yes | yes | yes | yes | Match, 2% early-payment discount |
| 104 | TAILSPIN LOGISTICS | €595.00 | TL-4401 | yes | no | yes | no | yes | Needs a person: tie |
| 104 | TAILSPIN LOGISTICS | €595.00 | TL-4402 | yes | no | yes | no | yes | Needs a person: tie |
| 105 | ACCOUNT FEE | €9.90 | — | — | — | — | — | — | No invoice: booked to bank fees |
| 106 | FABRIKAM OFFICE SUPPLY | €1,190.00 | NW-2026-0815 | yes | no | no | no | no | Rejected: the amount trap |
| 106 | FABRIKAM OFFICE SUPPLY | €1,190.00 | NW-2026-0901 | yes | no | no | no | no | Rejected: the amount trap |
| 106 | FABRIKAM OFFICE SUPPLY | €1,190.00 | 2026/0930 | no | no | yes | yes | yes | Match, collective payment |
| 106 | FABRIKAM OFFICE SUPPLY | €1,190.00 | 2026/0935 | no | no | yes | yes | yes | Match, collective payment |
Every row above is one query over the live bank feed and the open invoices, in ScramDB, computing five deterministic signals: exact amount, a 2% early-payment discount, the supplier's bank account, the invoice number buried in the payment reference, and the supplier's name in the counterparty. Plain SQL plus ScramDB's built-in full-text functions. No external service.
And the rows are where it gets interesting, because rules alone get this wrong:
- Payment 106 is a trap. It is exactly 1,190.00 euros, and there are two open Northwind Cloud invoices of exactly 1,190.00. An amount-matcher books it to Northwind and moves on. But the bank account is Fabrikam's, the name is Fabrikam's, and the payment reference lists two Fabrikam invoice numbers. It is one transfer settling two invoices: 357.00 plus 833.00. A batch payment.
- Payment 102 came through a payment app. No bank account match, no invoice number, and the counterparty reads
WOODGROVE PAY *CONTOSOADS. ScramDB's tokenizer, correctly, does not thinkCONTOSOADSis the wordcontoso. Only the amount agrees. This is judgment, not arithmetic. - Payment 103 is short by exactly 24.00 euros. That is not a partial payment, it is a 2% early-payment discount taken inside the discount window, and it needs its own booking.
- Payment 104 is a genuine tie. Two open invoices from the same supplier, same amount, same bank account, and a payment reference that says only "Freight August". There is no signal that separates them. Any system that picks one is guessing.
So the signals narrow the field, and the AI decides. Claude reads each candidate set from ScramDB, makes a call, and writes it down with a confidence and a reason in plain accountant's English:
| Outcome | Payments |
|---|---|
| Matched automatically | 3 |
| Booked directly, no invoice | 1 |
| Needs a person | 2 |
| Payment | Decision | Confidence |
|---|---|---|
| 101 | NW-2026-0815 | 0.99 |
| 102 | CA-77812 via Woodgrove Pay | 0.87 |
| 103 | 2026/0921, discount | 0.98 |
| 104 | TL-4401 or TL-4402 | 0.50 |
| 105 | bank fee, booked to fees | 0.98 |
| 106 | 2026/0930 + 2026/0935 | 0.97 |
Four of six payments resolved without a human: three matches, including the early-payment discount with a booking suggestion to the discounts account, plus the bank's own account fee, which has no invoice at all and gets booked straight to bank fees. Two go to a person: the payment-app transfer because nothing but the amount confirms it, and the Tailspin tie because nothing could.
That is the whole point. Not "the AI does everything". The AI does everything it can justify, and it knows exactly which two it cannot.
An append-only audit trail
Here is where most AI accounting demos stop, and where your customer's auditor starts asking questions.
Auditors and tax authorities everywhere ask two things of digital bookkeeping: that every entry can be traced and understood, and that a recorded entry cannot be quietly changed. They do not care whether a human or a model proposed a booking. They care that the decision, and every correction to it, is recorded and cannot be rewritten.
ScramDB gives you the database half of that out of the box, with nothing but ordinary SQL:
-- The matcher can read the inputs and append decisions. That is all.
CREATE ROLE matcher_agent;
GRANT SELECT ON invoices, bank_tx TO matcher_agent;
GRANT SELECT, INSERT ON match_decisions TO matcher_agent;
-- And nobody, not the agent, not the table owner, can edit history.
CREATE TRIGGER match_decisions_no_rewrite
BEFORE UPDATE OR DELETE ON match_decisions
FOR EACH ROW EXECUTE FUNCTION match_decisions_append_only();
When the agent tries to raise its own confidence on a decision it already made:
ERROR: permission denied: role "matcher_agent" requires UPDATE on table
"public.match_decisions" column(s): confidence
When the table owner tries the same thing:
ERROR: match_decisions is append-only: record a correction as a new row
So when the controller approves the payment-app match, she does not edit the AI's row. She appends her own, with her own role on it. The log reads top to bottom as exactly what happened: the AI proposed, a person confirmed, and both are on the record forever. Add ScramDB's continuous WAL archiving and point-in-time recovery, and you can reconstruct the state of the books at any moment you are asked about.
Compliance covers the whole process, including the process documentation your customers keep. ScramDB delivers the database's part of it out of the box: an append-only decision log that carries the reason and the actor, and the ability to rewind to any moment.
Validation on a branch of production
You would not let a new junior accountant post to the live books on day one. Do not ship a new matcher straight to your customers' books either.
CREATE DATABASE matcher_dry_run CLONE scramdb;
That is an instant, full, real copy of production in ScramDB. It shares storage with the source instead of copying it, so it is instant regardless of size. Point the agent at matcher_dry_run, let it decide a whole month of payments, compare its decisions with what was actually booked, then DROP DATABASE and nothing it did ever touched the real ledger.
Reconciliation status on demand
The decisions are just rows in ScramDB, which means they are data the Semantic AI layer can govern like any other. Model a view over payments and their latest decision, and the same conversation that answered the CFO now answers the controller:
"How much is still unmatched for Litware, and what is waiting on me?"
One query call. Two open invoices, 1,785.00 euros between them, and neither is due yet. Two decisions waiting for a person. No export, no dashboard refresh, no "let me pull that for you tomorrow".
The case for a single database
Look at what those two workflows actually touched: a general ledger, a budget, a live bank feed, e-invoice data, an AI decision log, and a semantic layer on top of all of it. In most stacks that is five systems and three pipelines for your team to build and babysit, and every pipeline is a place where the numbers stop agreeing.
In ScramDB it is one live copy. The CFO's question and the matcher's decision read the same rows, the instant they are committed. That is what I mean by UTAP, and finance is where the difference stops being architecture and starts being money.
And the data never leaves your infrastructure. Run ScramDB in your own data center or your own cloud region. No third party sees your customers' bank statements on the way to being matched, which, for anyone who has ever negotiated a data processing agreement, is not a small thing.
For the team building the product:
- Semantic AI comes with every ScramDB: the semantic layer, the hints and the instructions. Stop building a metrics layer, a vector store, an audit log and a sync job. ScramDB is all four, on one copy.
- Your agents program against a governed surface instead of raw tables, so the demo that works is also the product that is correct.
- The e-invoicing wave is about to hand your customers a mountain of structured invoices. Be the product that matches them.
For the CFOs using it:
- Ask the ledger anything, in a sentence, and get the number their own finance team defined.
- Every answer shows its SQL. "Where did this come from?" stops being a meeting.
- A new question does not wait for a new dashboard.
For the accounting firm with four hundred clients:
- Row-level security per client, enforced per request. One ScramDB, every client isolated.
- The matcher clears the obvious majority and hands over only the cases that need a professional.
- Every decision, human or AI, is on an append-only record you can walk an auditor through.
For the security review your customers will run:
- Agents run with least privilege and cannot rewrite what they did.
- Filters are data, so there is nothing to inject through the semantic surface.
- The semantic server runs sealed, inside ScramDB, with nothing it was not granted.
Getting started
Get ScramDB running and Semantic AI is already on. Point Claude, or the agent in your own product, at http://127.0.0.1:9191/mcp, and ask the ledger something your customers have always wanted to ask it. The quick start takes a few minutes, and the Semantic AI docs cover every tool.
If you are building finance software, I would genuinely love to hear what you are matching, reconciling or forecasting. That is exactly the work ScramDB was built for.
And much thanks for actually reading it.

