Datatrail
BUYER'S GUIDE - UPDATED JULY 2026

Data Catalog Tools: The Best Data Catalog Software Compared for 2026

Fourteen catalogs checked against primary documentation, not marketing pages. Which ones really have column-level lineage, what each costs to run, and the three acquisitions and one dead project that most comparison lists still miss.

Jump to the table
Read-only No card to start
Lineage map
Lineage mapped from query history. Read-only connection.
0

Read-only connection. Datatrail never moves or mutates your data.

In short

Data catalog tools inventory every table, column, and dashboard in your data estate and attach the context needed to use them: meaning, ownership, lineage, and usage. The leading options in 2026 are Collibra, Alation, Atlan, Informatica from Salesforce, Microsoft Purview, data.world, Coalesce Catalog, Secoda, Select Star, Datatrail, and the open-source OpenMetadata and DataHub. They split into enterprise governance platforms, modern automated catalogs, and self-hosted open source. Column-level lineage is no longer the differentiator it was, since almost all of them now have it. What actually differs is how much the tool derives automatically versus how much your team curates by hand, and whether the vendor will show you a price.

Last updated July 2026

// COMPARE

Side by side

Data catalog software compared

Tool Best for Lineage granularity Deployment Pricing
Datatrail Warehouse teams that want a catalog derived from lineage, not maintained by hand Column-level, native Read-only connection Published, self-serve
Collibra Regulated enterprises running a formal governance program Column-level SaaS or self-hosted No published pricing
Alation Analytics-first organizations that want strong search and adoption Table-level default, column-level is an add-on SaaS or customer-managed No published pricing
Atlan Modern data teams wanting fast adoption and native column lineage Column-level, native Single-tenant SaaS plus agent No published pricing
Informatica from Salesforce Large estates already standardized on Informatica integration Column-level SaaS plus Secure Agent No published pricing
Microsoft Purview Azure-centric organizations that want catalog and compliance together Column-level, with gaps SaaS plus integration runtime Published rates
data.world Teams that want a knowledge-graph catalog inside ServiceNow Column-level SaaS only No published pricing
Coalesce Catalog Teams that want catalog and transformation in one tool Column-level, native SaaS, pushes compute to warehouse Published: free tier, then $150 per user per month
Secoda Business teams discovering data inside Atlassian tools Table and column-level SaaS No published pricing
Select Star Snowflake shops wanting automated catalog and usage analytics Column-level, shown in a side panel SaaS No published pricing
OpenMetadata Engineering teams that want a full open-source catalog and will run it Column-level, native Self-hosted, four services Open source, managed tiers unpriced
DataHub Engineering teams wanting open-source metadata at scale Column-level, in the open-source edition Self-hosted, four dependencies Open source, cloud unpriced
Apache Atlas Existing Hadoop estates that already run the stack Column-level for Hive only Self-hosted on Hadoop services Open source
Amundsen Nothing new. See the maintenance warning below No automatic column-level extraction Self-hosted, five services Open source

Capabilities reflect each vendor's own documentation as checked on 31 July 2026. Only figures a vendor publishes itself appear anywhere on this page. Vendors move quickly in this category, so confirm current details directly before you buy.

// BASICS

The job to be done

What a data catalog is actually for

Strip away the category language and a data catalog exists to answer two questions that otherwise cost a data team an enormous amount of time. Where is the right data for this question, and can I trust this number. Everything a catalog ships, search, glossaries, ownership, lineage, popularity, quality badges, is in service of one of those two.

That framing matters at purchase time, because the failure mode of this category is well documented and always the same. A team buys a catalog, runs a discovery project, populates a few thousand entries with descriptions, and eighteen months later half of those descriptions refer to columns that no longer exist. The catalog becomes another artifact nobody trusts, which is the exact problem it was bought to fix. Manual curation does not survive contact with a warehouse that changes weekly.

So the question worth asking every vendor is not whether they support a business glossary. They all do. It is how much of the catalog is derived automatically from what the warehouse already knows, and how much depends on a human remembering to update it. Automated technical metadata, harvested schemas, query-derived popularity, and lineage parsed from query history stay correct on their own. Business context, ownership, and certification do not, and that is the part your program has to be designed around.

Lineage is where the two halves meet. A catalog entry that lists a column is a fact. A catalog entry that also shows the eleven models and two executive dashboards reading that column is a decision you can act on, which is why we build the catalog out of the lineage graph rather than the other way round. If you want that side in depth, the data lineage tools guide covers it, and data lineage vs data catalog works through where the two categories genuinely differ.

// 14 OPTIONS

The tools

What each data catalog tool is actually good at

Datatrail

Column-level, native

Datatrail comes at the catalog from the lineage side. It connects to Snowflake, BigQuery, Redshift, Databricks, or Postgres with a read-only role, parses query history alongside your dbt manifest, and the catalog falls out of the graph it builds: every table and column listed with its upstream sources, downstream consumers, and how often it is actually queried. Nothing is curated by hand, so nothing goes stale the week after the rollout. It is deliberately not a governance suite. If you need policy workflows, stewardship approvals, and regulatory reporting, buy one of the platforms below instead.

Collibra

Column-level

Collibra is the reference implementation of enterprise data governance: business glossaries, stewardship workflows, policy management, and a Gartner Leader placement in the 2026 Magic Quadrant for Data and Analytics Governance Platforms. In 2026 it has moved off catalog language almost entirely and now positions as an enterprise AI control plane, adding an AI Command Center in May 2026 on top of its Raito and Deasy Labs acquisitions. Two practical notes: column-level lineage is not generated for tables created by SQL statements unless you supply the SQL through the folder-connection method, and the CLI lineage harvester reached end of life on 31 July 2026, so connectivity now runs through Edge sites.

Alation

Table-level default, column-level is an add-on

Alation built its reputation on discovery and adoption: behavioral analysis of query logs to surface which tables people actually trust, plus a genuinely good search experience. In July 2026 it relaunched as AIOS, an intelligence operating system for enterprise AI, following its Numbers Station acquisition. The detail buyers most often miss is lineage granularity. Alation documentation states that table-level lineage is the default and that column-level lineage requires per-connector enablement, and for Snowflake specifically it requires a paid parser add-on. If column-level lineage is why you are buying a catalog, get that quoted explicitly rather than assuming it is included.

Atlan

Column-level, native

Atlan is the strongest of the modern catalogs on developer experience and time to value, with native column-level lineage that is core rather than an add-on, backed by dedicated column-process entity types and native OpenLineage ingestion from Airflow, Spark, and dbt Cloud. In 2026 it positions as the context layer for AI. Architecture is a single-tenant SaaS control plane in a dedicated VPC plus a Kubernetes Secure Agent running inside your environment to extract metadata locally, so it is not self-hostable end to end. Pricing is a conversation, not a page.

Informatica from Salesforce

Column-level

Salesforce completed its acquisition of Informatica on 18 November 2025 in a deal valued around 8 billion dollars, and the catalog survived intact: it still ships as Cloud Data Governance and Catalog inside IDMC, now branded Informatica from Salesforce. The 2026 direction is folding Agent Fabric and CDGC into a unified context catalog for AI agents, with new scanners for Salesforce Data 360 and Agentforce Marketing. Pricing runs on consumption units called IPUs, and while the model is documented, Informatica publishes no dollar rate per IPU. Hybrid and on-prem sources need a customer-installed Secure Agent sized at 16GB of RAM minimum.

Microsoft Purview

Column-level, with gaps

Purview Unified Catalog is the natural default if your estate is Microsoft, because governance, cataloging, and compliance share one control plane. It supports column-level lineage, though automated lineage is limited to a fixed list of process systems, and Microsoft documents that manual column-level lineage is not supported when a process asset sits between two data assets. One thing to be careful about: the classic Purview Data Catalog is in customer support mode and closed to new customers, but it has no published retirement date, and it is a different product from Azure Data Catalog, which was retired in May 2024. Microsoft is one of the few vendors that publishes real numbers.

data.world

Column-level

data.world built its catalog on a knowledge graph rather than a relational metadata store, which is what lets it connect business concepts to column-level technical lineage through the feature it calls Eureka Explorer. ServiceNow acquired the company in 2025, closing in July, and it now also surfaces as Data Catalog inside ServiceNow Workflow Data Fabric. That is either the main reason to buy it or the main reason not to, depending on whether your organization runs ServiceNow. Deployment is SaaS only, and pricing is a contact form.

Coalesce Catalog

Column-level, native

This is the product formerly sold as CastorDoc. Coalesce acquired it in March 2025 and folded it in as Coalesce Catalog, so a comparison that still lists CastorDoc as an independent vendor is out of date. The proposition now is governance embedded in transformation rather than a standalone catalog, with column-level lineage and pre-change blast-radius analysis across Snowflake, Databricks, Fabric, and BigQuery, plus Redshift in private preview. Coalesce is one of the very few vendors here that publishes real prices: a free Developer tier and a Starter plan at 150 dollars per user per month, billed against a shared action counter. The catalog is bundled rather than priced separately.

Secoda

Table and column-level

Secoda is an AI-native discovery layer aimed at business users as much as engineers, with real table and column lineage, though BI column lineage covers Looker and Tableau only. Atlassian acquired it in December 2025 and is folding it toward Rovo and the System of Work, which means data discovery lands where non-technical people already are, inside Jira and Confluence. If your problem is that analysts and operators cannot find or trust data, that distribution is a genuine advantage. Pricing is sales-led across Core, Premium, and Enterprise, with no published figures and no free tier.

Select Star

Column-level, shown in a side panel

Select Star automates almost all of the cataloging work, deriving documentation, popularity, and column-level lineage from query logs with very little manual input, which is why small data teams like it. Snowflake announced its acquisition in November 2025 and is folding it into Horizon Catalog. That is worth thinking about before you sign: if you are a Snowflake shop it is now a first-party direction, and if you are multi-warehouse it is a lock-in question. Note also that the visual lineage graph is object-level, with column-level relationships shown in a side panel rather than drawn on the graph.

OpenMetadata

Column-level, native

OpenMetadata is the most complete open-source catalog available and it genuinely has column-level lineage, so do not believe comparisons that claim otherwise. Apache 2.0, built by the team behind Hadoop, Atlas, and Uber Databook. The honest cost is operational: a production install runs a Java server, a MySQL or Postgres database, an OpenSearch or Elasticsearch cluster, and an ingestion runner, roughly 10 vCPU and 40GiB of memory before your first user logs in. Breaking changes have shipped in patch releases, and Snowflake lineage ingestion wants ACCOUNTADMIN-class privileges. Collate, the commercial sponsor, publishes tiers and limits but no dollar figures.

DataHub

Column-level, in the open-source edition

DataHub is the other serious open-source option, originally from LinkedIn. Acryl Data renamed itself to DataHub in May 2025 alongside a 35 million dollar Series B, and the open-source edition is branded DataHub Core. Column-level lineage is in the free edition, not gated behind cloud, currently across 53 of its connectors, with Kafka, object storage, and most BI-adjacent sources table-level only. Running it means Kafka, a relational database, a search index, and a graph index, plus four DataHub services. Release cadence is healthy, with 1.6.0 shipping in May 2026. DataHub Cloud publishes no pricing at all; the pricing URL returns a 404.

Apache Atlas

Column-level for Hive only

Atlas was the governance backbone of the Hadoop era and it still works, but it was designed for a world of Hive, HBase, Kafka, and Solr rather than cloud warehouses. There are no native hooks for Snowflake, BigQuery, or Databricks, its column-level lineage is essentially Hive-only, and standing it up means running the supporting cluster services. If you already operate Hadoop it is free and reasonable. If you are cloud-native, choosing it in 2026 means building the connectors nobody else will build for you.

Amundsen

No automatic column-level extraction

Amundsen came out of Lyft and shows up on almost every best-data-catalog list published this year. Those lists are out of date. The repository has had zero commits to its main branch in the last twelve months, the last commit landed on 2 April 2025, the last tagged release was August 2024, the last published Docker images date to March 2024, and 58 pull requests sit open and unmerged. Its reference deployment still pins Neo4j 3.5, which has been end of life since 2021. As of July 2026 the amundsen.io domain no longer serves the project. It also never had automatic column-level lineage extraction, only a data model and an off-by-default UI that deep-links to an external tool. We would not start a new deployment on it.

// CONSOLIDATION

Read this before shortlisting

Half this category was acquired in the last eighteen months

Data cataloging consolidated hard, and most comparison articles still present these as independent companies. If you are building a shortlist, the ownership question is now a real part of due diligence, because it determines roadmap direction and it determines whether the tool becomes a wedge into a much larger platform contract.

  • Informatica joined Salesforce on 18 November 2025, in a deal valued around 8 billion dollars. The catalog continues as Cloud Data Governance and Catalog and is being folded into Salesforce's agent governance story.
  • data.world joined ServiceNow, announced in May 2025 and closed that July. It now also appears as Data Catalog inside ServiceNow Workflow Data Fabric.
  • Secoda joined Atlassian in December 2025, heading toward Rovo and the Atlassian System of Work.
  • Select Star joined Snowflake, announced November 2025, folding into Horizon Catalog.
  • CastorDoc joined Coalesce in March 2025 and is now sold as Coalesce Catalog. The CastorDoc brand no longer exists.

None of that is automatically bad. ServiceNow and Atlassian distribution genuinely puts data discovery in front of people who would never open a catalog, and Snowflake owning Select Star is a plus if you are a Snowflake shop. The point is to price it in. Ask directly how the standalone product is funded post-acquisition, whether support for competing warehouses is a roadmap commitment or a legacy obligation, and what happens at renewal when the parent platform wants a bundle.

// OPEN SOURCE

Check the commit log

The open-source data catalog situation, checked rather than repeated

Open-source catalogs are a legitimate choice and two of them are genuinely good. The problem is that the published lists recommending them are copied from each other and are years stale, so here is what the repositories actually looked like on 31 July 2026.

OpenMetadata and DataHub are both healthy and both have real column-level lineage. OpenMetadata shipped through 2026 with an active release cadence, and DataHub tagged 1.6.0 in May 2026 with more than a hundred commits in the previous month. Column-level lineage is in DataHub's free Core edition rather than gated behind the paid cloud, currently covering 53 of its connectors. Either is a defensible foundation.

Amundsen is not. This is the one worth saying plainly, because it still appears on nearly every best-data-catalog list published this year. The main branch has had zero commits in the last twelve months and zero in 2026. The last merged pull request was on 2 April 2025, and it was an administrative change moving a maintainer to emeritus status. The last GitHub release was in August 2024, the last published Docker images including the latest tag date to March 2024, and 58 pull requests are open with none merged in roughly sixteen months. The reference deployment still pins Neo4j 3.5, which reached end of life in 2021. As of July 2026 the amundsen.io domain no longer serves the project or its documentation. There is also a widely repeated claim that Amundsen has column-level lineage: what exists is a data model plus an optional, off-by-default UI that deep-links to an external tool, with no parser that extracts column lineage from any source.

The honest accounting on the two healthy options is operational cost. OpenMetadata in production wants a Java server, a relational database, a search cluster, and an ingestion runner, roughly 10 vCPU and 40GiB of memory before anyone logs in. DataHub wants Kafka, a database, a search index, and a graph index, plus four of its own services. Neither has a license fee, and both have an engineer's salary attached. That trade is fine if you have platform engineers with capacity, and it is a poor trade if the catalog project is one person's side quest.

// NATIVE

Before you buy anything

Your warehouse already ships a catalog

Every major platform now includes cataloging and lineage. For a single-platform team this is often enough, and it is always the right baseline to measure a purchase against. Here is what each one really covers, including the limits that tend to surface after rollout rather than before.

Databricks Unity Catalog

The most complete of the four. Unity Catalog captures lineage automatically down to the column level for queries run on Databricks and aggregates it across every workspace attached to the metastore, and it is on by default for workspaces created after November 2023. Two corrections worth making, because most comparison pages get them wrong: lineage in Catalog Explorer is retained indefinitely rather than for 90 days, with the system tables holding a rolling one-year window, and the long list of exclusions those pages cite has largely been removed from the documentation. Real remaining limits: nothing before 1 September 2024, no lineage preserved through a rename of a catalog, schema, table, view, or column, no column lineage when the source or target is referenced by path instead of table name, and RDDs are not captured. There is no separate Unity Catalog charge, but it requires the Premium plan or above.

Snowflake Horizon Catalog

Horizon is best understood as an umbrella brand over features that are documented individually rather than a product you install: lineage, tagging, classification, masking policies, data metric functions, and the Trust Center. Column-level lineage went generally available in the 9.3 release in February 2025 and requires Enterprise Edition or higher, with one year of retention and nothing before November 2024. Horizon itself is not a billable line item, though components such as data quality monitoring and classification meter serverless compute. Keep three names apart, because they get conflated constantly: Horizon is the governance catalog, Open Catalog is the managed Apache Polaris service for Iceberg and is now closed to new customers, and Openflow is NiFi-based ingestion and not a catalog at all.

AWS Glue Data Catalog and SageMaker Catalog

These are two different layers and the difference matters. The Glue Data Catalog is a metadata store: it holds databases, tables, and schemas for Athena, EMR, and Redshift Spectrum, and despite a vague line in its own documentation it has no lineage UI, no lineage API, and no column-level lineage. Column-level lineage lives in Amazon SageMaker Catalog, the layer built on Amazon DataZone, which does advertise automated column-level lineage. Read its support matrix before you rely on it: automation covers Glue crawlers and Redshift by default, while Glue console jobs and EMR need explicit configuration and Visual ETL in fine-grained mode is unsupported. DataZone was not renamed and still exists separately, though its release notes stop in 2024.

Google Knowledge Catalog

Renamed from Dataplex Universal Catalog on 10 April 2026, with the API, CLI, and IAM names left unchanged, and the pricing SKUs still carrying the Dataplex name. The legacy Data Catalog service began a phased shutdown on 1 June 2026. Column-level lineage is real but narrow: it covers BigQuery, and Dataproc since May 2026, while Dataflow, Composer, Vertex AI, Looker, and anything arriving through OpenLineage stay table-level. The documented limits are specific and worth checking against your workload, including no column lineage for load jobs or routines, none upstream of external tables, top-level columns only with no nested STRUCT or JSON, and a silent fallback to table-level lineage once a job creates more than 1,500 column links.

The shared limitation is the platform boundary. Each of these is excellent inside its own walls and blind the moment your data crosses into a second warehouse, an ingestion tool, or the BI layer where people actually read numbers. If your entire estate is one platform, start here and buy nothing. If a dashboard in Looker or Power BI reads a Snowflake table populated from a Postgres source, no native catalog will show you that whole path, and that gap is the honest reason to add a third-party tool. We cover the specific platform paths in Snowflake lineage, Databricks lineage, BigQuery lineage, and Redshift lineage.

// PRICING

Who will show you a number

Data catalog pricing, and the vendors who publish it

Most of this category will not show you a price without a call. Below is every vendor on this page that publishes real figures on its own site or price API, quoted as published and cross-checked on 31 July 2026. Collibra, Alation, Atlan, Informatica, data.world, Secoda, Select Star, and DataHub Cloud publish nothing.

Vendor Published pricing What to watch
Datatrail Starter $99, Team $299, Scale $799 per month Flat monthly plans, self-serve signup
Coalesce Catalog Free Developer tier, Starter $150 per user per month Catalog bundled with transformation, billed by actions
Microsoft Purview $0.0165 per governed asset per day, plus $15, $60 or $240 per data governance processing unit Assets only count once linked to a governance concept
AWS Glue Data Catalog $1 per 100,000 objects per month, $1 per 1,000,000 requests, $0.44 per DPU-hour for crawlers First million objects and million requests are free
Amazon SageMaker Catalog $0.40 per GiB-month of metadata, $10 per 100,000 requests, $1.776 per compute unit No per-user charge
Google Knowledge Catalog $0.089 per DCU-hour for lineage and quality, $2.00 per GiB-month of metadata storage The 100 DCU-hour free tier excludes the premium lineage SKU

A note on the numbers you will find elsewhere. Review and procurement sites publish confident-looking figures for Collibra, Alation, and the rest, and they contradict each other badly enough that at least some of them are wrong. One widely syndicated listing renders Collibra at a per-user monthly rate that is almost certainly an annual contract figure mangled into the wrong unit. We do not reprint any of them here, and neither should the comparison you build internally. Get the quote.

The pattern worth noticing is that the hyperscalers publish exact metered rates while the specialist catalog vendors publish nothing. That is a reasonable proxy for how each expects to be bought. Metered rates mean you can model cost from your own asset counts before talking to anyone. Quote-only means budget approval and a procurement cycle before you can even compare. Ours are on the pricing page for the same reason.

// 4 QUESTIONS

How to choose

Four questions that decide which data catalog tool you need

01

How much is automatic?

Ask what a brand new table looks like in the catalog on the day it is created, before anyone touches it. If the answer is an empty entry waiting for a description, you have bought a documentation project rather than a catalog. This single question predicts whether the tool is still accurate in two years better than any feature list.

02

Is column lineage included or extra?

Almost every vendor now claims column-level lineage, and the claims are not equivalent. Alation documents table-level as the default with column-level requiring per-connector enablement and, for Snowflake, a paid add-on. Others derive it natively. Get the granularity, the connector coverage, and the price in writing.

03

Where does coverage stop?

Native platform catalogs stop at the platform. Open-source connectors stop where the community stopped building. Ask specifically whether coverage reaches the raw landing tables upstream and the dashboards downstream, because those two ends are where trust questions actually get asked.

04

Who owns the vendor now?

Five of the products here changed hands in the last eighteen months. Ask how the standalone product is resourced post-acquisition, whether support for competing warehouses is a commitment or a legacy obligation, and what the renewal looks like when the parent wants to sell you a platform bundle.

// HONEST

Where we fit

When Datatrail is the right pick, and when it is not

Datatrail inverts the usual order. Rather than building a catalog and adding lineage to it, it connects to Snowflake, BigQuery, Redshift, Databricks, or Postgres with a read-only role, parses query history alongside your dbt manifest into a column-level graph, and lets the catalog fall out of that graph. Every table and column appears automatically with its upstream sources, downstream consumers, and real query frequency attached. Because it is derived rather than curated, it describes the warehouse as it is today, not as it was during the rollout.

That design has a clear boundary and it is worth being direct about it. We are not a governance suite. If you need stewardship workflows, policy management, certification chains, and audit reporting for a regulated program, Collibra and Informatica have spent fifteen years building exactly that and we have not. If your priority is getting business users to adopt a catalog inside the tools they already use, Secoda and Atlan have invested far more in that experience. If you want a free catalog and have engineers to run it, OpenMetadata is a real answer.

Where we are hard to beat is the case most warehouse teams are actually in. Far more in the warehouse than dbt models, a handful of dashboards executives genuinely read, engineers who cannot safely drop a column because nothing shows what reads it, and no appetite for a six-month governance rollout. For that team the catalog is a by-product and the useful part is column-level lineage plus downstream impact analysis before a change ships. See also the data observability tools and data quality tools guides for the monitoring side, or compare us directly against Select Star and Alation.

// FAQ

Questions people ask

Data catalog tools, answered

What is a data catalog?

A data catalog is a searchable inventory of an organization tables, columns, models, and dashboards, with the context needed to use them: what each field means, who owns it, where it came from, and what depends on it. It answers the two questions that waste the most analyst time, which is where do I find the right data and can I trust this number. Modern catalogs build most of that inventory automatically from warehouse metadata rather than from a spreadsheet somebody maintains.

What is a data catalog tool?

A data catalog tool is software that connects to your warehouses, pipelines, and BI tools, harvests their metadata, and turns it into a searchable catalog with lineage, ownership, and usage attached. The category splits three ways: enterprise governance platforms such as Collibra and Informatica, modern automated catalogs such as Atlan and Datatrail, and open-source projects such as OpenMetadata and DataHub that you host yourself. They differ far more in how much is automated than in what they claim to store.

Why do I need a data catalog?

You need one when nobody can answer where a number came from without asking a specific person. The concrete symptoms are analysts rebuilding metrics that already exist, engineers afraid to drop a column because they cannot see what reads it, and audits that take weeks of manual evidence gathering. A catalog is worth buying at the point where tribal knowledge stops scaling, which for most teams is somewhere between five and fifteen people touching the warehouse.

What is the best data catalog software?

There is no single best data catalog software, because the category serves different jobs. For a regulated enterprise governance program, Collibra or Informatica. For modern teams that want fast adoption and native column-level lineage, Atlan. For a catalog derived automatically from lineage with published pricing, Datatrail. For an open-source deployment you will operate yourself, OpenMetadata or DataHub. Match the tool to who maintains the entries, because that is what actually decides whether the catalog stays accurate.

What is the best open source data catalog?

OpenMetadata and DataHub are the two credible open-source data catalogs in 2026, and both have genuine column-level lineage at no license cost. Apache Atlas remains viable only for existing Hadoop estates. Avoid Amundsen for new deployments: its main branch has had no commits in twelve months and its last release was August 2024. The real cost of the open-source options is operational, since both leaders need a database, a search index, and several services running before anyone logs in.

Does Snowflake have a data catalog?

Yes. Snowflake Horizon Catalog provides discovery, tagging, classification, access policies, and column-level lineage natively, with lineage requiring Enterprise Edition or higher and retaining one year of history. Horizon is not billed as a separate product, though some components meter serverless compute. Its boundary is Snowflake itself, so anything that happens in your BI tool, your ingestion layer, or a second warehouse sits outside it, which is the usual reason teams add a third-party catalog on top.

Does Databricks have a data catalog?

Yes. Unity Catalog is the governance and catalog layer for Databricks, and it is enabled by default for workspaces created after November 2023. It captures column-level lineage automatically for queries run on Databricks and aggregates it across all workspaces on the metastore, retained indefinitely in Catalog Explorer. There is no separate charge, though it requires the Premium plan or above. Like Horizon, it stops at the platform boundary and does not follow your data into other warehouses or BI tools.

What is the difference between a data catalog and data lineage?

A data catalog is the inventory, telling you what exists and what it means. Data lineage is the map of relationships between those entries, telling you where each one came from and what it feeds. A catalog without lineage lists assets you still cannot safely change. Lineage without a catalog gives you a graph nobody can search. Most serious tools now ship both, and the useful question is which one the product derives automatically and which one it expects your team to maintain.

How much do data catalog tools cost?

Most catalog vendors publish no pricing at all, including Collibra, Alation, Atlan, Informatica, data.world, Secoda, and Select Star, and quote per deployment against connected sources, catalogued assets, and seats. The exceptions publish real numbers: Coalesce lists a free tier and $150 per user per month, Microsoft charges $0.0165 per governed asset per day plus processing units, AWS and Google publish per-request and per-DCU rates, and Datatrail publishes flat monthly plans. Treat third-party price estimates on review sites with suspicion, because they frequently contradict each other.

A catalog that builds itself from your lineage

Connect your warehouse read-only and get every table and column inventoried automatically, with upstream sources, downstream consumers, and real usage attached. Published pricing, no sales call.