
BI Architecture
Data integration for BI
Plan how source changes reach BI reports, keep figures traceable and define when a reporting period is complete.
Data integration for business intelligence (BI) moves source records into reporting under agreed rules. Start with one report: identify the records it needs, how changes reach its reporting dataset, when that dataset is complete, and who can explain a figure that looks wrong.
Trace one figure from source to report
For each important figure, record the source system, relevant keys, eligible records and extraction cutoff. Follow the records through preparation, the reporting dataset and the final report. Note where identifiers, dates or statuses change.
A service report might combine cases with team assignments. Its owner needs to decide which case states count, which team receives a case that moved, and how an unmatched assignment appears. Copying both tables successfully does not settle those rules.
| Stage | Decision to record | Evidence to keep |
|---|---|---|
| Source | Which records and changes are authoritative? | Source owner, keys and cutoff |
| Movement | How do new, corrected and deleted records arrive? | Run or stream identifier and processing status |
| Preparation | Which mappings and exclusions apply? | Transformation version and rejected-record details |
| Delivery | Which period can readers use? | Source coverage and successful report update |
These are responsibilities, not a requirement for separate products.
Azure Data Factory and Synapse pipeline monitoring exposes pipeline and activity runs, including copy input, output and errors, with duration and status. Copy activity details can include volume read and written, files or rows copied, throughput, applied configuration and the duration of execution steps. Keep the relevant run evidence with the release record so an unusual total can be investigated against what the copy processed.
Azure Data Factory Copy Activity metrics
- Files or rows copied
- Available in copy activity monitoring
- Throughput
- Measured during execution
- Duration and status
- Includes run and activity details
- Fault tolerance option
- Abort or continue with skipped incompatible data
Choose a movement pattern
A scheduled full extract can suit a modest, stable report. An incremental load moves selected changes but needs a dependable checkpoint and a way to recover a missed interval. Change data capture can convey inserts, updates and deletes where the source and service support it.
A near-real-time path may reduce transport delay, but preparation and the reporting layer still affect what readers see.
Choose the pattern against the reader's decision time and the source's capabilities. A checked daily load may serve a weekly review; an intraday queue view may need more frequent updates. Neither schedule makes an incomplete source period final.
CDC is an option only where the source and service support it. For example, Azure Data Factory’s documented CDC workflow uses SQL Change Data Capture information from Azure SQL Managed Instance or SQL Server to identify inserted, updated and deleted data for incremental movement.
Agree how those events alter the destination before treating an incremental run as equivalent to a complete period.
Keep the result traceable
Retain enough source identity and timing information to explain a reporting row. Define how retries, late corrections and deletions affect it. A repeated load should have a known write rule: replace, merge or append. Otherwise recovery can change a total without an explanation.
Separate job status from report completeness. A job can finish when an expected file is absent or records were excluded.
In supported Azure Data Factory Copy activity scenarios, fault tolerance can either abort the copy or continue while skipping incompatible data; session logging can record skipped data. Check the expected deliveries and applicable run details before releasing a figure.
Release a defined period
For each report, specify the latest source event it should include, the latest complete business period and what readers see if either is missed. In Power BI Import mode, the semantic model imports and holds copied data until a source data refresh brings in changes; an upstream load alone does not update that copy. Other connection modes have different behaviour.
A release rule can require expected source deliveries, completed transformations, selected checks for a fixed period and an updated reporting model. If a dependency fails, hold the affected result or retain the last accepted one with its earlier cutoff visible. Name the source, pipeline and report owners so corrections and changed fields have a clear response path.
DirectQuery and Direct Lake models, and live connections to Analysis Services, do not import data in the same way: they query the underlying source with user interaction. Record the model’s mode in the release contract so owners know whether a release depends on a refresh of copied data or on the state of the queried source.
Include privacy in the movement contract
Where reporting data includes personal information, make its movement and use part of the integration agreement. The Australian Privacy Principles (APPs) govern collection, use and disclosure, as well as organisational governance and accountability, information integrity and correction, and individuals’ access rights.
The APPs apply to organisations and agencies covered by the Privacy Act 1988. They are principles-based and technology neutral, so document how the reporting path handles relevant information rather than treating a particular transfer tool as the compliance rule.
In this guide
- Comparing scheduled batch loads with near-real-time pipelinesCompare batch and near-real-time paths by decision timing, report completeness, recovery and visible data delay.
- Defining data freshness expectations for a reportSpecify a BI report's source cutoff, ready-by time, completeness rule and reader-facing status when data is late.
- Handling missing source data without silently filling gapsDistinguish absent source deliveries from genuine zero activity and decide how a BI report should show incomplete data.
- Testing a pipeline after a source schema changeCheck extraction, mappings, delivered fields and report figures after a source schema changes.



