Data Platform Foundations¶
The foundations are a review system for data-platform decisions. They are not products, maturity badges, or a universal reference architecture.
Use them to ask three things before approving any change:
- Which foundation does this decision strengthen?
- Which foundation could it weaken?
- What evidence would prove the architectural claim?
| Foundation | Objective | Warning sign | Diagnostic question |
|---|---|---|---|
| Trust | Ensure data can support consequential decisions. | Consumers manually verify numbers before using them. | Would I bet an important business decision on this data? |
| Ownership & Accountability | Give every dataset, pipeline and product a clear owner. | “Ask the one person who understands it.” | Who fixes this tomorrow if it breaks? |
| Semantic Clarity | Keep business meaning consistent across the platform. | Two products report different values for the same metric. | Is this a data problem or a definition problem? |
| Reliability & Operability | Make processing predictable, recoverable and supportable. | Recovery depends on manual intervention or tribal knowledge. | Can another engineer troubleshoot this at 3 AM? |
| Observability & Transparency | Explain what failed, why and who is affected. | Consumers discover incidents before the platform team. | Who is affected, and how severely? |
| Cost Awareness | Make platform cost visible, attributable and justified. | Cloud spend grows without a measurable owner or outcome. | Is the value greater than the operating cost? |
| Safe Speed | Enable fast change without putting production trust at risk. | Experiments run directly against production assets. | How can this be tested without breaking production? |
| Evolution & Deletion | Let assets change, deprecate and disappear deliberately. | Everything is preserved forever “just in case.” | What happens if we delete this tomorrow? |
| Organizational Fit | Match the platform to the team's capacity and operating model. | The architecture requires expertise the team cannot sustain. | Does this scale with the team, not only the data? |
| AI Readiness | Make data and AI outputs governed, traceable and explainable. | AI work begins before quality and semantic consistency exist. | Can we explain this output end to end? |
How the foundations apply across layers¶
Every foundation applies across the full lifecycle, but its concrete implementation changes by layer.
| Layer | Primary responsibility | Foundation emphasis | Minimum evidence |
|---|---|---|---|
| Raw / landing | Preserve source data at the platform boundary through controlled ingestion, auditability and replay. | Trust, ownership, reliability, observability | Source contract, reconciliation, rejection behaviour, replay and latency |
| Curated / conformed / Silver | Create consistent and reusable business entities for downstream teams. | Semantics, trust, reliability, ownership, cost | Contracts, quality results, deterministic rebuild, change ownership and run cost |
| Business-ready / Gold | Publish governed data products and interfaces for BI, APIs, applications and AI. | Ownership, semantics, observability, trust, cost | SLO, consumer acceptance, lineage, access, usage and attributable cost |
External sources are outside the platform boundary. They remain visible as context, while contracts define the interface the platform can depend on. Ingestion is a capability within Raw / landing, not a separate layer. Layer names describe responsibilities rather than prescribing a physical implementation.
Architecture validation checklist¶
Before approving a pipeline, dataset, dashboard, data product or AI feature, verify:
- The business outcome and owner are explicit.
- The strengthened and weakened foundations are named.
- Behaviour is understood at 10× data volume and 10× consumers.
- Failure, recovery and replay can be demonstrated.
- Data meaning can be explained end to end.
- Cost is attributable to an outcome.
- The team can operate the design without hero dependency.
- Retention, evolution and deletion are deliberate.
The goal is not to build pipelines. The goal is to build a trusted, observable, governed, operable, cost-aware and AI-ready data platform.