Skip to content
NORDGLADE

Nordglade / Experiment 03

If someone challenged a number your AI helped prepare, what could you show about how it got there?

ArtifactAIOS Decision Record
Recorded2026-08-10
Length50 s
OutcomeCompleted

The recording

Paused at 44 s
Screen recording50 s · Fictional company data

Nordvind Energy is a fictional company created for this demonstration. The recording shows an AI assistant working from the records that company has saved, answering whether it has to report under CBAM this quarter. What you see here is what the assistant leaves behind afterwards. It shows where it looked, that its answer was held to those records, that two figures in it were traced back, and which company supplied the language model that wrote the wording. The quotation at the bottom is that language model’s own account of its work. It sits in a box of its own because it is a claim rather than a check. See it full size.

Proof and claim

What the software did is written down as fact. What the language model says it did is written down as a claim.

The five rows come from the software itself, so they hold whatever the answer happens to say. The sentence in the box underneath comes from the language model, describing its own work. As an explanation it is worth reading. As evidence it is worth nothing, and the record says so on its face rather than leaving you to work it out.

The record

Four questions
01

What it used

This is the same AI assistant as experiment 01, answering the same question about Nordvind Energy and CBAM. The record’s top row shows the one tool it reached for, search_context, which reads the company’s own stored information, and confirms it ran.

The wording of the answer was produced by a language model, deepseek-v4-flash, supplied by Lyceum Technology from Germany. That is worth one line: experiment 01 closes by admitting it ignored the software’s own advice about which language model suits questions with exact figures in them. This one followed it.

02

What it checked

Three checks ran. The record says what each one returned.

The row labelled grounding reads passed. It means the answer was built on the company’s own records, not on whatever the language model happens to know about CBAM. Where the assistant does reach past the records, it says so. One paragraph on screen is marked “my own read, not from your records.”

The row labelled number check reads ok · 2 figures verified. Two numbers in the answer were followed back to the records they came from. The chat’s own summary of the same exchange agrees: 2 figures trace to the tools.

The row labelled sourcing reads n/a. That check looks for the assistant saying “your records show” about a rule your records do not mention. It found no sentence of that shape to test. So it returned nothing at all.

What a passing check does not tell you

Two numbers came out of your records. That is not a statement that they are right, or that CBAM applies to your company. And nothing to test is not the same as a pass. That is why the record prints n/a rather than a tick.

03

Where it stopped

The answer states that the records hold no CBAM reporting threshold or tonnage limit. It marks its own general knowledge of CBAM as “my own read, not from your records.” It closes: “That decision is yours to make. Confirm it with your accountant or advisor.”

The record does not stop anything. It is written after the answer, and nothing in it blocked a question or changed a word. What it adds is that the point where the assistant handed the decision back is now on paper, not only on screen.

04

What this does not prove

A complete record does not make the answer correct. The number check covers numbers and nothing else, so an invented claim with no figure in it passes untouched. The language model’s quoted reason is a sentence it wrote, not a window into how it worked.

This is one question, on a fictional company, on my own machine. It does not show that a record gets written every time, or that these rows stay honest against a real company’s data.

What I take from it

Author's note

If a regulator, a customer or a bank questions a figure your team prepared with AI, the useful thing is not a better answer. It is being able to show what the software actually did, without reopening the source files and rebuilding the work from scratch. This entry is one question’s worth of that. The part I would defend is the labelling. The software’s own record and the language model’s account of itself are kept apart, because only one of them is evidence.

Recording limits

Two things visible in the recording that this page does not claim
BN-04The record is a local file, opened in a second tab

The browser’s tabs and address bar stay visible and are not retouched. The address bar shows the file path the card was opened from on my own machine. The card is an export view of the log, not a panel inside the product.

BN-05One line about Germany is not a claim about the whole system

The bottom row of the record names the company that supplied the language model for this answer, and the country it operates from. It says nothing about where the rest of the software runs, or where the company’s records are held.

Company
Fictional demonstration subject. No connection to a real company, client or filing.
Same artifact as