Most enterprises already have significant investments in their data platforms. Warehouses are deployed, pipelines are running, and dashboards are everywhere. The infrastructure, by most measures, is in place. And yet many AI initiatives still struggle to move from proof-of-concept to production.
The issue is rarely the absence of data. It is the absence of curated data. What organisations have built, in most cases, is data volume — years of transactions, events, logs, and records distributed across systems that were designed to capture information, not to make it consistently trustworthy across teams and use cases.
What they have not built, with the same rigour, is the practice of treating that data as something that needs to be continuously validated, classified, documented, and governed before it earns the right to inform a decision or train a model.
Industry research reflects this reality.
A decision or train a model.
Industry research reflects this reality.
Precisely’s Data Integrity Trends Report found that 64% of organizations name data quality as their single biggest data challenge.
Informatica’s CDO Insights survey of 600 data leaders found that 92% are concerned AI pilots are moving forward before the underlying data problems are solved.
That is the context for this blog. This blog is an honest look at what data curation requires, how Snowflake approaches it, where it genuinely helps, and where the work still falls on your team regardless of the platform.
Data governance challenges grew 89% between 2023 and 2025 according to the same report. You read that right. Governance problems nearly doubled in two years, while AI investment tripled. That gap is where most data initiatives go to die.
"FOMO-driven AI is exactly how companies end up with endless use cases and proofs of concept that never see production. Your goal isn't to use AI for the sake of AI. Your goal is to solve real business problems."
— Chandra Donelson, Chief Data & AI Officer — CDO Magazine Global Data Leadership Summit, 2025
“FOMO-driven AI is exactly how companies end up with endless use cases and proofs of concept that never see production. Your goal isn’t to use AI for the sake of AI. Your goal is to solve real business problems.”
— Chandra Donelson, Chief Data & AI Officer — CDO Magazine Global Data Leadership Summit, 2025
What Data Curation Actually Means
Data curation is often misunderstood as a one-time setup task.
In practice, it is an ongoing discipline that ensures data remains usable and trustworthy as systems evolve.
A curated data environment typically includes six core capabilities:
- Quality monitoring – measuring freshness, completeness, and integrity over time
- Metadata management – ensuring data is understandable and discoverable
- Access governance – enabling secure but broad usability
- Lineage tracking – tracing outputs back to their original sources
- Data discovery – helping users find the right assets quickly
- Ingestion discipline – defining clear quality expectations from the moment data enters the platform
When these elements work together, data moves from “we have it somewhere” to “we can trust it.” That transition is what unlocks enterprise-scale AI.
Why Snowflake Has Become Central to Data Curation
One reason Snowflake has gained traction in modern data architectures is its design philosophy.
Governance and data management capabilities are built into the platform rather than layered on top.
This means security, lineage, discovery, and monitoring share the same metadata foundation.
Several capabilities stand out.
Snowflake Horizon Catalog
Horizon is the unified governance and discovery layer inside Snowflake. It applies sensitive data classification, access policies, and lineage tracking consistently across every query engine that touches your data — Snowflake, Spark, any Iceberg-compatible tool.
The practical curation capabilities inside Horizon are worth naming specifically. Data Metric Functions run quality checks on a schedule: freshness, completeness, null rate, referential integrity. Column-level lineage in Snowsight traces data from source to final table, including ML feature views and model objects. Universal Search, powered by natural language, lets anyone find data assets without knowing the table name or schema.
At Snowflake BUILD 2025 in November, Horizon received semantic search updates built on NLP and RAG.
Data Metric Functions for In-Built Quality
One of the most practical tools available in Snowflake is Data Metric Functions (DMFs). DMFs allow teams to define SQL-based quality checks directly on data assets and run them continuously.
Examples include:
- freshness thresholds
- null rate monitoring
- uniqueness validation
- referential integrity checks
Instead of a one-time audit, teams gain continuous quality signals that can be monitored and acted upon.
This transforms data quality from a periodic review into an operational metric.
Dynamic Data Masking and Row-Level Security
Data Curation is fundamentally about the safe and controlled usability of data across the enterprise. In Snowflake, dynamic data masking operates at query time rather than at storage time, which means the original data remains intact while access is intelligently governed based on user roles and permissions.
There is no data duplication, no parallel masked environment to maintain, and no additional operational burden. This is what transforms self-service analytics from a theoretical aspiration into a secure, scalable enterprise capability.
Apache Iceberg and Open Format Governance
Snowflake’s commitment to Apache Iceberg — reaching full general availability in 2024 and extended at Summit 2025 with Catalog-Linked Databases. This means governance metadata lives with the data, not inside a proprietary system. Tags applied in Horizon travel with Iceberg tables regardless of which compute engine accesses them. For curation specifically, this solves a real engineering problem: consistency of policy enforcement across a heterogeneous compute environment.
Where Technology Alone Is Not Enough
Even the best data platform cannot replace disciplined ownership.
Most organizations already have some form of catalog, governance framework, or quality checks in place.
The challenge is consistency.
New datasets arrive without clear documentation or context, making them difficult to understand and use effectively. Quality checks are created once during initial setup but are rarely reviewed or updated as data and systems evolve. Over time, metadata becomes outdated or incomplete, leaving teams without reliable visibility into their data landscape.
Data curation only works when it is treated as a continuous operational practice, not a project milestone.
The Shift Enterprises Are Beginning to Make
The most successful AI programs today are reframing the problem.
Instead of asking: “Which AI use case should we build next?”
They start with:“Is our data environment ready to support AI at scale?”
That shift changes the priorities of the organization.
They start with:“Is our data environment ready to support AI at scale?”
That shift changes the priorities of the organization.
It leads to investments in:
- metadata and discovery
- automated quality monitoring
- governance embedded into the platform
- shared data definitions across teams
These capabilities may not appear in AI demos, but they determine whether AI systems survive beyond the pilot stage.
The Bottom Line
AI adoption is accelerating across every industry.
But AI systems are only as reliable as the data they learn from.
But AI systems are only as reliable as the data they learn from.


