AI × Finance · Week 1 · Topic 1.1

How AI Actually Works

The one question that protects every number you publish.

About 15 minutes · read + the exercise

Let’s open with a magic trick, the dangerous kind. Ask an AI for a bank’s gross NPA with no document in front of it and it answers in a heartbeat: ‘around 1.5%’. Calm, fluent, sure. The real number is 1.23%, sitting on page 4 of the deck. It missed, and it didn’t blink. Now here’s the part that should worry you. When it’s right, it uses the exact same calm, fluent, sure voice. AI has no ‘I’m not certain’ tone. A fact and a fabrication come out identical. This whole lesson is how you tell them apart, and it’s the wall the rest of the cohort is built on.

AI says a fact and a fabrication in the same confident voice. It can’t supply the doubt. Supplying it is your job, and it’s the entire skill.

Why it never says ‘I don’t know’

Start with the strangest thing about it: it will guess before it will admit ignorance, every time. Here’s why, and it’s not what people assume.

The exam that trained it
Picture a school exam where a blank scores zero, but a wild guess might fluke a mark. What do you do? You guess every question. Leaving one blank is the single move that can never help you. AI was trained on exactly that exam. In training, a confident guess scored better than ‘I don’t know’, so abstaining got optimised out of it. (OpenAI’s own 2025 research showed this plainly.) The confidence isn’t a personality or a bug. It’s a student who learned that a blank always scores zero, so it never leaves one.

What it’s actually doing, one level under ‘it’s smart’

When you type a question, the model does exactly one thing: it predicts the next word, then the next, from patterns it absorbed in training. No look-up. No database. No little auditor inside checking facts against a source. It is an extraordinary guesser of ‘what word probably comes next’, and that is the entire engine.

Everyday
Your phone’s autocomplete. Type ‘running late, reaching in’ and it offers ‘five minutes’. It doesn’t know your traffic. It completes a pattern it has seen a million times. AI is that, far smarter.
Finance
Ask for a bank’s NPA and it completes a plausible bank-NPA number the same way, ‘around 1.5%’. It didn’t look it up. It pattern-completed it. Which is fine for a guess and fatal for a published note.

The one question that protects you: CREATE or RECALL?

This is the whole skill, and it takes one second. Before you trust any answer, ask which mode you’re in.

CREATE (trust it, then skim)
Draft, rewrite, summarise, explain
Reformat a table, brainstorm
The pattern IS the answer
It shines here
RECALL (never trust it alone)
A specific number, date or quote
‘What did the filing say?’
No pattern to match, so it invents one
Make it cite the line, then you check
Same tool, same five seconds. The only thing that changed is create versus recall, and that’s the difference between a clean note and a correction under your name.

The trap: it writes like a pro and cites like a fraud

Here’s the move that catches careful people. Language has strong patterns, grammar, the rhythm of a research note, so AI writes beautifully. But one number from one bank’s one quarter is a fact to look up, and it cannot look anything up. So it drafts a flawless paragraph and slips an invented figure inside, in the same elegant prose. The better it writes, the more you trust it, and the more dangerous that one wrong number becomes. Fluent does not mean correct. On a research desk, the polish is the disguise.

Watch it lie, then watch it cite

Pure recall vs sourced recall · Axis gross NPA
How you askedWhat you got
No document: ‘what’s Axis’s gross NPA?’‘Around 1.5%’: confident, fluent, wrong
Deck attached: ‘quote the exact line’1.23%, with the source: asset-quality chart, p4
It wasn’t lying the first time, it pattern-completed a plausible number. Pure recall failed; sourced recall worked. That gap is the entire reason this discipline exists.
The trap careful people fall into
Ask AI to ‘summarise the key numbers’ from a results PDF and you’ve wrapped a create verb around a recall task. It hands back a gorgeous paragraph with one quietly invented figure inside. It looks perfect, which is exactly why it gets used. When a ‘summary’ contains specific numbers, it’s secretly a recall task. Treat it like one, and make every number cite its line.
My one strong opinion here
Create versus recall is the load-bearing wall of this whole cohort. Ninety percent of ‘AI for finance’ advice is noise. Get this one distinction and everything downstream, prompting, verification, the lot, is just plumbing. Miss it and no clever prompt saves you, because you’ll trust the wrong outputs with total confidence. If you take one thing from these four weeks, take this.

Catch it lying, right now

Five minutes, hands on the keyboard
Pick a company you cover. (1) Ask its last-quarter NIM with no document. Note how sure it sounds. (2) Attach the real deck and ask again, adding ‘quote the exact sentence proving the number’. (3) Compare. The first is usually subtly wrong; only the sourced one is right. You just watched the same confident voice do both, on a name you know cold.

Create, or recall?

Tap each task.0 / 5
‘Rewrite this note paragraph to be tighter.’
‘What was SBI’s gross NPA last quarter?’
‘Explain how NIM is calculated.’
‘Pull the credit cost from this deck.’
‘Summarise the key numbers from this PDF.’
The rule: create is drafting, explaining, reformatting, trust then skim. Recall is any specific fact, never trust it alone. A ‘summary’ that contains numbers is recall wearing a create costume.
War story
Early on I asked an AI for a bank’s net interest margin and it gave me a clean 3.4 percent. Confident, instant, wrong. No filing, it just predicted the kind of number that usually follows that question. That was the day it clicked: this thing finishes sentences, it doesn’t look things up. Treat it like autocomplete with a silver tongue and you stop trusting it and start feeding it.

When it breaks, and the fix

It gave a confident number that turned out wrong.
Pure recall. Attach the source and make it quote the exact line. No quote, no number.
A ‘summary’ slipped in a figure that isn’t in the deck.
A recall task hid inside a create one. Treat any summary with numbers as recall.
The better it read, the more you trusted it.
Fluent is not correct. On anything you publish, the polish should make you check harder, not less.
It said ‘approximately’ and you relaxed.
‘Approximately’ is still a guess. For published work, get the cited line.

Quick questions people actually ask

So I can’t trust it with numbers at all?
You can, once they’re sourced. Untrusted recall plus a cited line you read equals a trusted number. The cite is the bridge.
Why does it sound so sure when it’s wrong?
Because in training, abstaining scored worse than guessing. Confidence is the trained default, not a signal that it knows.
Is this just a Claude thing?
No. Every large language model works this way. The create-versus-recall question protects you on all of them.
Starter pack
Your Week 1 deliverable: a one-page Trust Map. Two columns, ‘I trust AI for this’ (create) and ‘I always verify this’ (recall). You’ll use it every day. Open the pack →

Do this week

The model can’t tell a fact from a fluent guess, they come out of the same machine, in the same voice. You can, in one question: did I ask it to create, or to recall? Everything else in this cohort is built on that one question.
Sources: OpenAI, ‘Why language models hallucinate’ (2025), the clearest official account of why models guess instead of abstaining. Axis gross NPA pulled from its Q4FY26 investor presentation via the QuarterBook store at build time.
/finqrate AI × Finance Cohort · Topic 1.1 · for members only, licensed to you personally.