Research library
ERPData

ERP Operational Intelligence: Why Text-to-SQL Fails Without Business Context

Text-to-SQL is not enough for ERP and operating data. Teams need business definitions, table context, visible SQL, and validation before answers can be trusted.

Article brief

Author
Dotnitron
Published
May 2, 2026
Read time
4 min read
ERP Operational Intelligence: Why Text-to-SQL Fails Without Business Context

Key takeaways

  • ERP questions fail when the model has SQL ability but not business context.
  • The hard part is selecting the right tables, joins, dates, definitions, and filters before SQL is written.
  • A governed answer layer should expose SQL, result tables, assumptions, and validation evidence.

Text-to-SQL demos look simple: ask a question, generate SQL, return a chart. ERP reality is not simple. A business question like why did revenue fall in the western region may involve orders, invoices, credits, customers, product hierarchies, territory mappings, currency rules, cancellations, and reporting definitions that differ by team.

The failure mode is dangerous because the answer can look professional while being wrong. The SQL may run. The chart may render. The narrative may sound confident. But if the wrong tables, joins, filters, or definitions were used, the business answer is not trustworthy.

The hard part happens before SQL generation

In complex ERP environments, the model's ability to write SQL is not the scarce capability. The scarce capability is context selection. Which tables matter? Which table represents booked revenue versus billed revenue? Which date field should be used? Which customer hierarchy is official? Which records should be excluded? Which team definition of margin applies?

That context usually lives across schema names, table descriptions, old analyst queries, BI model logic, finance definitions, operational playbooks, and tribal knowledge. A governed answer layer has to retrieve and assemble that context before the model writes SQL.

What a governed ERP answer should include

  • The business interpretation of the question, including synonyms and likely definitions.
  • The selected tables, joins, filters, date fields, and metrics used to answer it.
  • The generated SQL, visible to business analysts and data owners.
  • The result table and chart, so the narrative is not the only output.
  • The assumptions and validation checks that should be reviewed before the answer is reused.

Why dashboards do not solve every question

BI dashboards are still valuable. They standardize recurring reporting and keep core metrics consistent. But every operating team eventually asks questions that are not already modeled: why did this segment move, which vendors changed behavior, which invoices are unusual, where did the control exception start, what changed since last month?

Those questions become tickets, analyst interruptions, spreadsheet workarounds, or delayed decisions. Dotnitron implements governed answer workflows for that follow-up layer, using capabilities such as SemeLabs when visible SQL, approved scopes, and validation evidence are the right pattern.

A safe first rollout

Start with one team and one approved data scope. Collect 25 to 50 real business questions. Generate answers with visible SQL. Have business and data owners score each answer as pass, partial, fail, or out of scope. Use the results to identify missing definitions, weak table context, and rollout boundaries.

The goal of the pilot is not to prove that AI can write SQL. The goal is to prove that a team can trust, review, and reuse answers without creating a new analyst bottleneck.

What an answer packet should show

A useful ERP answer packet should show the business question, interpreted meaning, approved data scope, selected tables, generated SQL, result table, chart if needed, assumptions, validation checks, and unresolved caveats. That gives the business owner and data owner a shared artifact to inspect instead of a black-box answer.

This is especially important in portfolio, finance, and operating environments where one wrong definition can change the conversation. Revenue may mean booked, billed, recognized, net of credits, or grouped by a management hierarchy. Margin may be calculated before or after allocations. Customer churn may depend on invoice inactivity, contract status, or usage. The system has to expose those choices.

Where governed natural-language analytics creates value

  • Executive follow-up questions after a dashboard review.
  • Recurring variance explanations where analysts repeatedly pull the same tables.
  • Portfolio-company operating reviews where data definitions differ across systems.
  • Finance and operations queues where teams need a traceable answer faster than a BI ticket cycle.

The implementation detail that matters most

The implementation should not let users query every table freely on day one. Start with approved semantic zones, business definitions, and question classes. Log every query, generated SQL, validation result, and user correction. Over time, those corrections become the operating memory that makes the answer layer safer and more useful.

Use this guide

Turn the article into a working session.

Pick one workflow from the article and map it against your own team. Write down the input sources, current manual steps, reviewer decisions, output format, and the metric that would prove the workflow is worth automating.

  • What work should agents prepare before a human reviews it?
  • Which documents, data sources, tools, or approved system connections would the workflow need?
  • What output would make a reviewer say, this saves real time?

Ready to turn one painful workflow into a working AI system?

Bring the workpaper, evidence review, ERP answer queue, diligence step, or reporting loop your team wants to stop doing manually.