Translate a vague request into an answerable analytical question
In this article (7 sections)
Translate a vague request by identifying the decision, defining the measure and population, specifying the comparison and checking whether the available data can answer it. A precise query is useful only after the business question has a precise meaning.
The process should reduce ambiguity without pretending that every request can be answered from the files already available. Sometimes the correct result is a narrower answer and a clear request for missing evidence.
Find the decision behind the wording
Suppose someone asks, “Why is revenue down?” Ask what decision they need to make: investigate a discrepancy, change a campaign, adjust inventory or explain a management report.
Then clarify what “revenue” means and which periods are being compared. Completed orders, invoices and cash collections can move differently. A fall in one measure is not automatically evidence that another fell.
Do not begin by selecting a chart or asking an assistant to produce possible reasons. First establish whether the reported decline exists under a consistent definition.
Turn the request into a sequence
A useful sequence is: verify the measure, establish the comparison, decompose the observed difference and investigate plausible mechanisms. Each step has a different evidence requirement.
For the original synthetic commerce data, an answerable first question is: “What is January completed-order amount under the current eligibility rule?” The result is 104,000 paise across eight orders.
The same files do not establish why a real business's revenue declined. They do not contain a comparable prior period for that full question or a causal design. The metric contract makes the narrower scope explicit.
Use a question specification
| Field | Example |
|---|---|
| Decision | Reconcile the January headline before using it in a review |
| Measure | Completed-order amount, not recognized revenue |
| Population | Eligible completed orders, including unmatched customers |
| Period | January 1 inclusive to February 1 exclusive, declared local timezone |
| Comparison | Source-grain reference versus the disputed report |
| Evidence | Orders, query logic, eligible IDs and source version |
| Output | Reconciled amount, discrepancy explanation and limitation |
This specification is short enough to review with a stakeholder and detailed enough to guide a query. It also helps an AI assistant identify what it may calculate and what it must not invent.
Distinguish descriptive and causal questions
“Which product group contributed most to the observed amount difference?” can be answered by a defined decomposition when comparable data exists. “Did the campaign cause the difference?” requires evidence about the intervention and an appropriate comparison design.
Do not change the wording from “associated with” to “caused by” merely because the stakeholder wants a decisive answer. Explain what additional evidence would make the causal question assessable.
Similarly, “How many orders appear in the refund ledger?” is answerable from the supplied records. “How much refund cash left the business in January?” is not answerable from an undated refund ledger.
Check feasibility before committing to an output
Inspect whether the required columns, time coverage, identifiers and definitions exist. Identify missingness and source-quality issues that could invalidate the requested comparison.
If the full question is not answerable, offer the supported part and state the remaining requirement. For example: “I can calculate completed-order amount and reconcile the two queries; I need dated collection records to analyze January cash received.”
That is more useful than either fabricating an answer or refusing the entire request without explaining what can be done.
Confirm the interpretation in writing
Summarize the question, contract, output and unresolved assumptions. If the stakeholder changes the decision, revisit the specification rather than simply adding another visual to the same analysis.
Keep the original wording as context, but make the agreed analytical question the basis for acceptance checks. The final answer should show how it addresses that question and where its scope ends.
Exercise: translate “Which customers are best?” into two distinct analytical questions. Define the measure and observation window for each, then explain why the resulting customer rankings might legitimately differ.
NeuraPath's Data Analytics with Generative AI course connects stakeholder needs with data preparation, querying and interpretation. An answerable question gives the technical work a clear purpose and an honest boundary.
Continue learning
This article is part of the Metrics, visualization and decision communication sequence. Use the neighbouring tasks when you need the prerequisite or the next application.
- Review the prerequisite or neighbouring task in Run a dashboard requirements interview.
- Continue with Explain why two correct reports can disagree.
Pankit Kumar has 10 years in Data Science & AI, building and shipping production systems in regulated pharma and clinical environments. He is a freelance trainer at Boston Institute of Analytics, AnalytixLabs and Scaler, and has taught this material to thousands of working professionals.
This article is part of our Data Analytics with Generative AI programme — 3–4 months. The full analyst stack — Excel, SQL, Power BI and Python pipelines — then a generative-AI layer you can prove is right.
Explore Data Analytics with Generative AI