Logo
Departments
Resources
Company
Contact Us
News
Executive Reporting

Natural language generation: how to tell whether a report a machine wrote can be trusted

September 17, 2026
www.predictx.com/resources/travel-analytics-who-performed-the-analysis
www.predictx.com/resources/kpi-reporting-purpose-and-audience
www.predictx.com/resources/natural-language-generation-report-trust
www.predictx.com/resources/executive-buy-in-travel-programme
www.predictx.com/resources/executive-dashboard-purpose-and-limits
www.predictx.com/resources/out-of-policy-travel-spend-behavioural-audit
www.predictx.com/resources/travel-disruption-modal-shift-analysis
www.predictx.com/resources/strategic-sourcing-corporate-travel
www.predictx.com/resources/t-e-leakage-detection-agentic-ai
www.predictx.com/resources/agentic-ai-corporate-travel-leakage-detection
www.predictx.com/resources/travel-leakage-supplier-negotiations
www.predictx.com/resources/how-to-reduce-corporate-travel-leakage
www.predictx.com/resources/how-to-measure-corporate-travel-leakage
www.predictx.com/resources/business-travel-emissions-reporting-scope-3-csrd
www.predictx.com/resources/corporate-travel-leakage-causes
www.predictx.com/resources/ai-travel-analytics-te-reporting-fails
www.predictx.com/resources/agentic-ai-corporate-travel-expense-management-guide
www.predictx.com/resources/entity-level-travel-spend-analytics
www.predictx.com/resources/hotel-attachment-rate-missing-spend-te
www.predictx.com/resources/vendor-negotiation-intelligence-corporate-travel
www.predictx.com/resources/te-policy-simulation-calculator-predictx
www.predictx.com/resources/business-travel-emissions-scope3-data-gaps
www.predictx.com/resources/travel-and-expense-data-analytics-the-agentic-ai-shift
www.predictx.com/resources/travel-and-expense-management-predictx-solutions
www.predictx.com/resources/corporate-travel-carbon-reporting-data-quality
www.predictx.com/resources/agentic-ai-travel-management-modern-edge-t-e
www.predictx.com/resources/predictx-global-conflict-module-corporate-travel-risk-management
www.predictx.com/resources/future-global-travel-expense-management-ai
www.predictx.com/resources/multi-modal-models-vs-ocr
www.predictx.com/resources/continuous-air-sourcing-travel-and-expense-data
www.predictx.com/resources/travel-data-predictive-analytics-net-zero-targets
www.predictx.com/resources/agentic-ai-air-sourcing-imperative
www.predictx.com/resources/audit-ready-reporting-data-chain-of-custody-tne
www.predictx.com/resources/predictx-ai-rubicon-te-reporting-sustainability
www.predictx.com/resources/whitepaper-modal-shift-co2e-savings-audit-ready-p1
www.predictx.com/resources/csrd-scope3-business-travel-emissions-compliance
www.predictx.com/resources/predictx-wins-major-business-travel-technology-innovation-award-t-ereporting
www.predictx.com/resources/btsa-2025-ai-commitment-gap-t-e-reporting-agentic-ai
www.predictx.com/resources/complete-guide-air-sourcing-navigator-agentic-ai-corporate-travel-managers-t-e
www.predictx.com/resources/continuous-air-sourcing-ai-solution-travel-managers-travel-expense
www.predictx.com/resources/audit-ready-esg-compliance-travel-emissions-data-product-sheet
www.predictx.com/resources/air-sourcing-navigator-agentic-ai-product-sheet
www.predictx.com/resources/predictx-in-focus-1-agentic-ai-auditable-sustainability-and-the-future-of-elite-t-e-reporting
www.predictx.com/resources/the-future-proof-travel-program-ensuring-agility-with-advanced-business-intelligence
www.predictx.com/resources/the-5-solution-ai-fluency-blueprint-cogent-agentic-ai
www.predictx.com/resources/ai-project-root-causes-strategic-data-cogent-agentic-ai
www.predictx.com/resources/ai-failure-mcdonalds-aircanada-t-e-reporting-cogent-agentic-ai
www.predictx.com/resources/keesup-choe-btn-interview-cogent-agentic-ai-predictx
www.predictx.com/resources/corporate-travel-sustainability-car-rental-emissions
www.predictx.com/resources/what-to-ask-your-ai-prompts-cogent-agentic-ai
www.predictx.com/resources/the-last-mile-corporate-carbon-footprint-ghg-compliance
www.predictx.com/resources/how-agentic-ai-powers-t-e-reporting-rag
www.predictx.com/resources/predictx-squake-master-last-mile-emissions-esg-compliance
www.predictx.com/resources/beyond-dashboards-the-agentic-ai-revolution-in-t-e-reporting-expense-audit
www.predictx.com/resources/2025-playbook-for-net-zero-business-travel-by-predictx-squake-actionable-sustainability
www.predictx.com/resources/t-e-manager-tomorrow-cogent-agentic-ai-corporate-travel
www.predictx.com/resources/predictx-squake-audit-ready-co2-reporting
www.predictx.com/resources/introducing-predictx-in-focus-newsletter-your-corporate-travel-data-advantage
www.predictx.com/resources/cogent-tech-hotlist-agentic-ai-travel-and-expense
www.predictx.com/resources/prompt-engineering-cogent-agentic-ai-guide-t-and-e
www.predictx.com/resources/keesup-choe-agentic-ai-cogent-travel-expense
www.predictx.com/resources/product-sheet-cogent-ai-powered-travel-and-expense-t-e-reporting
www.predictx.com/resources/6-powerful-cogent-use-cases-for-t-e-reporting-travel-data
www.predictx.com/resources/cogent-wins-2025-bts-europe-innovation-faceoff
www.predictx.com/resources/transform-t-e-management-with-cogents-ai-powered-solutions
www.predictx.com/resources/predictx-squake-sustainability-travel-management
www.predictx.com/resources/unlocking-net-zero-goals-in-corporate-travel-insights-from-predictx-at-itm-sustainability-showcase-2024-webinar
www.predictx.com/resources/accelerating-towards-net-zero-a-guide-for-corporate-travel-managers
www.predictx.com/resources/unveiling-predictxs-internal-carbon-pricing-tool-a-transformative-leap-for-business-travel-sustainability
www.predictx.com/resources/how-ai-can-help-track-and-reduce-your-companys-travel-emissions
www.predictx.com/resources/defining-success-the-north-star-metric-for-business-travel-managers
www.predictx.com/resources/enhancing-corporate-travel-management-with-predictx-scorecard-a-comprehensive-solution
www.predictx.com/resources/understanding-compliance-and-risk-in-corporate-travel
www.predictx.com/resources/internship-journey-at-predictx-a-blend-of-learning-growth-and-inspiration
www.predictx.com/resources/unlocking-the-power-of-data-in-corporate-travel-discover-the-story-by-predictx
www.predictx.com/resources/effective-meetings-and-events-management-in-corporate-travel
www.predictx.com/resources/optimizing-corporate-card-usage
www.predictx.com/resources/adhering-to-the-csrd-shaping-the-future-of-corporate-travel-sustainability-with-predictx
www.predictx.com/resources/optimizing-spend-amidst-record-corporate-travel-industry-growth-in-2024
www.predictx.com/resources/the-simulation-engine-showcases-at-btn-us-innovate-2021
www.predictx.com/resources/how-to-promote-sustainability-and-calculate-your-companys-carbon-footprint
www.predictx.com/resources/spend-less-time-managing-cross-border-activities
www.predictx.com/resources/stay-on-top-of-tax-compliance
www.predictx.com/resources/keep-up-to-date-with-pre-trip-data
www.predictx.com/resources/manage-employee-generated-spend
www.predictx.com/resources/scorecard-for-travel-and-expense-management
www.predictx.com/resources/tickets-refunds-and-asset-recovery
www.predictx.com/resources/spend-reporting-made-simple
www.predictx.com/resources/sourcing-and-policy-management
www.predictx.com/resources/simulate-your-budget-with-predictx
www.predictx.com/resources/from-reactive-reporting-to-proactive-management
www.predictx.com/resources/hollywood-studio-transforms-the-analysis-process
www.predictx.com/resources/enhancing-travel-management-how-predictx-transformed-savings-analysis-for-a-european-retailer
www.predictx.com/resources/transforming-travel-data-enhancing-quality-and-efficiency-with-predictx
www.predictx.com/resources/predictx-for-susutainability
www.predictx.com/resources/predictx-for-egs
www.predictx.com/resources/predictx-for-travel
www.predictx.com/resources/predictx-uses-machine-learning-to-power-business-reporting
www.predictx.com/resources/merging-fragmented-travel-data-into-one-dynamic-system
www.predictx.com/resources/enhancing-traveler-satisfaction-and-employee-wellbeing-in-corporate-travel

Natural language generation is the part of a reporting system that turns structured data into written sentences. It does not decide what the data means. In a well-built system a separate deterministic step finds the result first, and natural language generation only puts that finding into readable English.

There are two questions you can ask about a report a machine wrote, and most people ask only the first. The first is whether the writing is any good. The second is whether the sentences correspond to the data.

Only the second matters, and it is settled long before the first word is generated, by a decision most vendors never put in writing: which part of the system decided what was worth saying.

Corporate travel buyers already know what it feels like to be handed a number they cannot check. Speaking to Business Travel Executive in May 2025 for its travel buyers' point-of-view special report, one corporate travel buyer, Mark Ziegler, put the position plainly: "We can't measure or audit dynamic rates. We are left to trust the GDS and the hotels."

A report written by a language model can repeat that position or it can end it, and nothing about how the report reads will tell you which. The failure mode has a name and a literature: the survey Ji and colleagues published in ACM Computing Surveys in 2023 catalogues it across summarisation, dialogue and data-to-text systems, and observes that generation of this kind is "prone to hallucinate unintended text".

That is not a reason to refuse machine-written reports. It is a reason to know which component was allowed to decide what the report says, and this article explains that distinction precisely enough to use on any vendor, including us.

In this article

  1. What is natural language generation?
  2. How does natural language generation work in reporting?
  3. What is the difference between an AI that finds a result and an AI that describes one?
  4. How would you know if an AI invented a figure?
  5. How do you audit an AI-generated report for accuracy?
  6. Why do two systems describe the same data differently?
  7. What should you check before sending an AI-written report to your CFO?
  8. Does an AI-generated report still need a human to review it?
  9. Frequently asked questions

What is natural language generation?

Natural language generation is the production of readable text from non-linguistic input such as a database, a set of metrics or a model output. It covers three separate jobs: deciding what to say, deciding how to word it, and producing grammatical sentences. Those jobs are distinct, and in reporting the first one carries almost all of the risk.

The field has been careful about that separation for a long time. Ehud Reiter and Robert Dale set out the standard pipeline in Building Natural Language Generation Systems, Cambridge University Press, 2000, and the shape they described is still how practitioners talk about it:

StageWhat it decidesThe reporting version
Document planningWhat to say, and in what orderWhich findings are worth a sentence this month, and which lead
MicroplanningHow to word it, what to group into one sentence, what to call each thingWhether this is "spend rose sharply" or "spend rose 14% against plan"
Surface realisationProducing grammatically correct textThe finished English

Read that table again with one thing in mind. The stage that decides whether a claim is true is the first one, not the last, so a system can be flawless at the last two and still tell you something that never happened.

This is why "the report reads well" is not evidence of anything. Fluency is produced at stage three. Truth is produced at stage one.

How does natural language generation work in reporting?

In reporting, natural language generation sits at the end of a pipeline that has already turned raw transactions into metrics. Two designs are common. In the first, the model reads the data and decides what is interesting. In the second, a deterministic layer decides what is interesting and the model is handed the finding to write up.

Both produce a page of prose about your programme. They are not the same product, they do not carry the same risk, and the difference is invisible from the finished page.

Take one sentence a reader might see in a monthly report on a travel programme:

Illustrative example, not a PredictX finding. Advance purchase compliance fell to 61% in August, down from 74% in July, driven mainly by two business units in the Northern Europe region.

In design one, a model was shown a table and asked what stood out. The percentages, the direction, the attribution to two business units and the word "mainly" were all produced by one component in one step, and the only record of how they were arrived at is the sentence itself.

In design two, a query ran at the warehouse. It compared August against July on a defined compliance measure, applied a threshold for what counts as a material fall, ranked the contributing units and returned a structured finding: measure, period, current value, prior value, delta, top contributors. The model received that object and wrote one sentence from it.

The sentences are identical. The second one can be traced to the query that produced it. The first one cannot be traced to anything.

Why the order matters more than the model

A model asked to find something will find something. That is a property of the task rather than a defect in any particular system: given a table and an instruction to report what is notable, a generative system returns a notable-sounding claim whether or not one exists in the data.

A model asked only to narrate a finding that already exists has a narrower job. It can still write a clumsy sentence or state a delta less clearly than a person would. It cannot produce a finding the detection step did not produce, because it was never asked to look.

That is the whole argument, and it is worth saying what it is not. Several platforms build this way and describe it in similar terms, so this is no claim that one vendor has solved something nobody else has. It is a way to read a product description, and a test you can apply to any of them.

A second question sits underneath this one, and it is about the party rather than the component: who produces the analysis, and what that organisation is paid to do. This article stops at the architecture, which is the half you can check from a product description.

What is the difference between an AI that finds a result and an AI that describes one?

A detector decides that something in the data is worth reporting. A renderer turns that decision into a sentence. When one component does both, a well-formed sentence is your only evidence that the finding is real. When they are separate, the finding exists as a record before any text is written, and the text can be checked against it.

The consequence is about evidence rather than about accuracy in the abstract.

If detection and narration are one step, "how do you know this is true?" can only be answered by re-reading the output or running it again to see whether it says the same thing. Neither is a check: the first inspects the claim using the claim.

If detection happens first and deterministically, the answer is a different kind of object: a rule, a threshold, a period, a query, and a row of numbers that either satisfies the rule or does not. You can disagree with the threshold. You can argue that 5% is the wrong bar for a material change. What you cannot do is be shown a finding that was never detected.

This is how PredictX builds Overture, its monthly reporting product for a corporate travel programme. Detection runs deterministically at the warehouse, and the language model narrates what detection found rather than searching for findings of its own. For the worked version rather than the principle, how the monthly travel briefing is written sets out the full sequence.

How would you know if an AI invented a figure?

From the finished page, usually you would not. Invented figures are not misspelled or oddly phrased; they are plausible, correctly formatted and consistent with the surrounding sentences. That is why the check cannot be a reading check. It has to be a structural one, made once about how the report is produced rather than every month about what it says.

The obvious response is to validate the output: read it, sample a few numbers against the source, sign it off. That catches arithmetic errors and misses the harder failure, a sentence whose numbers are all correct and whose claim is not.

"Driven mainly by two business units" is a causal claim. Confirming that the units exist and that their figures are right tells you nothing about whether they drove anything.

Validation after the fact also scales badly in the direction reporting scales. A monthly report across twelve business units and six categories holds too many claims to check one at a time, which is why it was automated.

So move the check earlier. If the claim came from a rule you can read, on data you can query, before any text existed, you audit the rule once and every month's output inherits it.

How do you audit an AI-generated report for accuracy?

Ask five questions about how the report was produced, not five questions about what it says. Between them they establish whether the finding existed before the sentence did. A vendor whose system detects first can answer all five in concrete terms. A vendor whose model decides what is significant will answer the first two in adjectives.

We call this the Detection Test: five questions that establish whether a report's sentences were written from findings, or written in order to find them. Use it on us. The point of publishing it is that it works on any vendor.

#Ask thisA sound answer sounds likeA weak answer sounds like
1Which component decided this sentence was worth writing?"A detection step in the data layer. The model received the finding already formed.""Our AI reviews your data and surfaces what matters."
2Show me the rule that fired, and its threshold.A named measure, a comparison period, a numeric threshold, and how contributors are ranked."The model assesses significance from context."
3When was this text written, and against which version of the data?A generation timestamp and a named data snapshot, both visible to the reader."It is written when you open the page."
4Where is the calculation method for each number in it?Definitions, calculation methods and exception lists available inside the product."That is in the documentation somewhere."
5What happens if I open the same report this afternoon?The same text, because it was written once against a fixed snapshot and stored."You may get slightly different wording."

Questions three and five look like small operational details. They are the two that most often separate a report you can quote in a meeting from one you cannot.

If narration is produced on demand, the words and the numbers are generated at different moments and can drift apart. Two people reading the same month can be shown two different sentences about it, and neither is wrong, which is worse than one being wrong. Live generation is a fair choice for an interactive assistant and a poor one for a document somebody will forward to a CFO.

The opposite choice is to write the narration once, at refresh, against a fixed monthly snapshot, and store it. Overture works this way deliberately: re-running the pipeline part-way through a month still surfaces the previous month's snapshot, so the sentences always describe the numbers printed beside them. It is not real time, and that constraint is the point rather than a gap.

Why do two systems describe the same data differently?

Because they disagree about what counts as significant, not usually about the arithmetic. Two systems given the same month of travel data will compute similar totals and then diverge on which movements deserve a sentence, how much of a change is worth mentioning, and what to attribute it to. Those three decisions are configuration, not intelligence.

Three places an AI reporting pipeline can diverge

The threshold: one system reports any month-on-month movement over 5%, another over 10%, a third uses a statistical test against the last twelve months. Same data, different number of findings, and none of them is wrong.

The denominator: compliance measured against bookable trips is a different number from compliance measured against all trips, and both are defensible. A written report that does not expose its denominator is asking you to accept a definition you have not seen.

The attribution: "driven mainly by" is the highest-risk phrase in any generated report. It needs a ranking rule and a cut-off, and where those are undefined the phrase is decoration attached to a real number.

Ask any vendor to state those three and the conversation moves from whose prose is better to whose definitions you agree with, which is the conversation worth having.

What should you check before sending an AI-written report to your CFO?

Three things, and none of them takes long. Check one number against the source system. Check that the period label on the text matches the period label on the figures. Then find the strongest causal claim in the document and ask whether the system that wrote it had a rule for making that claim.

The first two take five minutes and catch plumbing errors: a stale snapshot, a mislabelled month, a currency conversion applied twice.

The third is a single-pass check rather than a monthly chore, and the trick is to read for the verbs. "Rose", "fell" and "was" are descriptions, carrying only the risk of the underlying figure. "Drove", "caused" and "driven mainly by" are claims about mechanism, and each one needs a rule behind it.

If the vendor can show you the rule, you never check that class of sentence again. If they cannot, you now know which sentences in every future report you are personally standing behind.

Does an AI-generated report still need a human to review it?

Yes, and detection-first architecture changes what the review is for rather than removing it. When a deterministic step has already established that a finding is real, the reviewer is no longer verifying arithmetic. They are deciding whether a real finding matters to this business this month, which is a judgement no pipeline makes.

That is a smaller job than checking a report and a more senior one. A system can correctly detect that hotel attachment fell nine points in one region and be entirely unaware that the region ran a three-week office closure. The finding is true, and the interpretation belongs to somebody who was there.

The industry is heading into this question rather than around it. Suzanne Neufang, chief executive of the Global Business Travel Association, framed the association's May 2026 research on technology and managed travel around "closing the gap between what is possible and what travel programs experience today". Generated narrative is one of the things now arriving in that gap, and the practical question for a programme is not whether to allow it but how to check it.

The argument here is that you check it once, where the findings are made, rather than every month where the sentences are made.

Frequently asked questions

Can you trust an AI-generated report?

You can trust it to the extent that the findings in it were established before the text was written. If a deterministic step detected the result and the model only put it into sentences, the report is as reliable as the underlying rules and data. If the model decided what was significant, the sentence is the only evidence.

How does a machine write a report from raw data?

Transactions are cleaned and joined into metrics, a detection step compares those metrics against thresholds and prior periods to produce structured findings, and natural language generation turns each finding into a sentence. Composition then arranges the sentences into a document. The writing is the last step, not the analytical one.

What is the difference between natural language generation and a large language model?

Natural language generation is the task of producing text from non-linguistic data, and it predates large language models by decades. A large language model is one way to perform part of that task, usually the wording and sentence-forming part. The decision about what to say can sit outside the model entirely.

Can an AI-generated report be audited after the fact?

Only if the system kept the finding separately from the sentence. Where detection produces a structured record with its measure, period, threshold and inputs, any sentence can be traced back to the record that produced it. Where the model produced both at once, there is nothing behind the text to audit.

Is an AI-written report accurate enough for a board or an executive?

Accuracy is the wrong test on its own, because a report can be arithmetically accurate and still assert a cause it never established. The test that matters at that level is traceability: whether each claim in the document can be walked back to a rule, a period and a set of numbers.

Can a report be written without anyone preparing it?

Yes, and that is the practical reason narrative reporting is automated at all. Where detection runs on a schedule and narration is generated at refresh and stored, the document exists before anyone asks for it. What cannot be automated is the decision about which real finding deserves action this month.

See the mechanism at work

This distinction is easier to judge against a worked example. Overture produces a written monthly briefing on a corporate travel programme: detection runs deterministically at the warehouse, and the model then picks the most salient of the findings it is handed and writes them up. It chooses what leads, never what counts as a finding, and it cannot reach past the set detection produced.

Show me how the month's findings become a briefing

Related Posts

September 3, 2026

What an executive dashboard is for, and the month it stops being opened

Dashboards confirm numbers you already chose. They cannot say which matters this month, so they go unopened.
September 9, 2026

Why leadership does not fund your travel programme, and what executive buy-in turns on

Leadership rarely funds a travel programme whose numbers come from the suppliers it pays. Know who produced each figure.
September 19, 2026

KPI reporting: what a KPI report is for, and why most of them go unread

A KPI report is for one named reader, not a circulation list. Most are assembled from what systems measure.
September 14, 2026

Travel analytics: who performed the analysis of your corporate travel programme?

Who performed your travel analysis, and what that organisation is paid to do, is part of the evidence.
No items found.