
Intro
More than ten years after “big data” was first coined, data remains one of the most vital and rapidly expanding engines of innovation across both large companies and fresh startups. From delivering health checks that are basic to business operations to intelligently automating routine work with machine learning, data has become the core nervous system for decision-making in organizations of every size. In addition, the use of data now extends far beyond data scientists, data analysts, and data engineers — everyone is both a producer and a consumer of data.
The consequence of this heightened attention to data: the business of managing data has already emerged as one of the fastest-growing segments of infrastructure, estimated at over $70B and representing more than one-fifth of all enterprise infrastructure spend in 2021. What makes this market’s formation especially compelling is that it brings together software engineering, analytics, and artificial intelligence, while also benefiting from the powerful momentum of cloud computing. (For more on the architectural evolution and the forces behind this massive trend, see this piece, Emerging Architectures for Modern Data Infrastructure, which was just updated for 2022.)
The expansion of the data sector has also created some of the most exciting and consequential enterprise software companies of the past several years. Recent public heavyweights such as Snowflake and Confluent have already transformed how thousands of companies run and how millions of products are assembled. Still, many people know less about the operators and builders — the next wave of companies defining their categories.
To help make sense of the noise after a record-setting 2021 in which data companies took in tens of billions in venture capital funding — and an already strong 2022 — we’ve put together the first class of the Data50. These are the leading companies across the most compelling categories in data. Collectively, these 50 companies are valued at more than $100B and have raised roughly $14.5B in total capital, with 20 having become unicorns by 2021.
Without further delay, we’re pleased to present the Data50 of 2022.
The Data50 list
Show
All categories
- All categories AI/ML BI & Notebooks Customer Data Analytics Data Governance & Security Data Observability ELT & Orchestration Query and Processing
Apply Filter Clear all filters
RankCompanyCategoryLocationValuation RangeWebsite
1 Databricks Query and Processing San Francisco, CA $5B+ Databricks 2 Fivetran ELT & Orchestration Oakland, CA $5B+ Fivetran 3 Scale.ai AI/ML Palo Alto, CA $5B+ Scale.ai 4 OneTrust Data Governance & Security Atlanta, GA $5B+ OneTrust 5 Dbt labs ELT & Orchestration Philadelphia, PA $1B-$5B Dbt labs 6 Starburst Query and Processing Boston, MA $1B-$5B Starburst 7 Collibra Data Governance & Security Brussels, Belgium $5B+ Collibra 8 Dremio Query and Processing Santa Clara, CA $1B-$5B Dremio 9 Dataiku Query and Processing New York, NY $1B-$5B Dataiku 10 Hugging Face AI/ML New York, NY $250-999M Hugging Face 11 DataRobot Query and Processing Boston, MA $5B+ DataRobot 12 Primer.ai AI/ML San Francisco, CA $250-999M Primer.ai 13 Snorkel AI/ML Palo Alto, CA $1B-$5B Snorkel 14 Anyscale AI/ML San Francisco, CA $1B-$5B Anyscale 15 Firebolt Query and Processing Tel Aviv, Israel $1B-$5B Firebolt 16 Astronomer ELT & Orchestration Cincinnati, OH $100-$249M Astronomer 17 Alation Data Governance & Security Redwood City, CA $1B-$5B Alation 18 Weights & Biases AI/ML San Francisco, CA $1B-$5B Weights & Biases 19 Sigma Computing BI & Notebooks San Francisco, CA $1B-$5B Sigma Computing 20 Monte Carlo Data Observability San Francisco, CA $250-999M Monte Carlo 21 OctoML AI/ML Seattle, WA $250-999M OctoML 22 Census Customer Data Analytics San Francisco, CA $250-999M Census 23 Hex BI & Notebooks San Francisco, CA $250-999M Hex 24 Hightouch Customer Data Analytics San Francisco, CA $250-999M Hightouch 25 Amperity Customer Data Analytics Seattle, WA $1B-$5B Amperity 26 BigID Data Governance & Security New York, NY $1B-$5B BigID 27 Privacera Data Governance & Security Fremont, CA $250-999M Privacera 28 Immuta Data Governance & Security Boston, MA $250-999M Immuta 29 Bigeye Data Observability San Francisco, CA $250-999M Bigeye 30 Matillion ELT & Orchestration Greater Manchester, United Kingdom $1B-$5B Matillion 31 Heap Customer Data Analytics San Francisco, CA $1B-$5B Heap 32 Tecton AI/ML San Francisco, CA $250-999M Tecton 33 Imply Query and Processing Burlingame, CA $250-999M Imply 34 Sisu Data BI & Notebooks San Francisco, CA $250-999M Sisu Data 35 Rudderstack ELT & Orchestration San Francisco, CA $100-$249M Rudderstack 36 ActionIQ Customer Data Analytics New York, NY $250-999M ActionIQ 37 ClickHouse Query and Processing Portola Valley, CA $1B-$5B ClickHouse 38 Airbyte ELT & Orchestration San Francisco, CA $1B-$5B Airbyte 39 Rockset Query and Processing San Mateo, CA $250-999M Rockset 40 Labelbox AI/ML San Francisco, CA $250-999M Labelbox 41 Explorium AI/ML San Mateo, CA $250-999M Explorium 42 Rasa AI/ML San Francisco, CA $100-$249M Rasa 43 Prefect ELT & Orchestration Washington, DC $250-999M Prefect 44 Materialize Query and Processing New York, NY $250-999M Materialize 45 Coiled AI/ML New York, NY $100-$249M Coiled 46 Preset BI & Notebooks San Mateo, CA $100-$249M Preset 47 Metabase BI & Notebooks San Francisco, CA $100-$249M Metabase 48 Iterative.ai AI/ML San Francisco, CA $100-$249M Iterative.ai 49 Robust Intelligence AI/ML San Francisco, CA $100-$249M Robust Intelligence 50 Fiddler AI/ML Mountain View, CA $100-$249M Fiddler
Methodology
Data50 companies were founded after 2008, have raised new funding in the last two years, and their employee base is growing at at least 30% YoY. Their products are horizontal technologies serving data or data application teams across industries.
Rankings are based on a blend of most recent valuation, company size, employee growth over the last two years, years in operation, and current revenue scale. Employee data is based on publicly available data from LinkedIn. Funding data is based on publicly available data from Pitchbook and Crunchbase, and is accurate as of March 22, 2022.
Note that this list does not include transactional database companies such as CockroachDB, PlanetScale, and Yugabyte because usage of the data with those technologies is inherently transactional instead of analytical.
Looking beneath the surface, we’ve divided the Data50 into seven subcategories.

- Query and processing technology is the central engine for accessing, aggregating, and computing data. It includes two primary classes: batch processing (e.g. Databricks and Starburst) and real-time processing (e.g. ClickHouse and Imply). The latter has drawn increasing attention over the past few years, fueled by rising demand for real-time applications. AI/ML (artificial intelligence and machine learning) includes software that uses algorithmic modeling and machine learning to process large-scale data. This area is maturing and thriving, as shown by the sheer number of companies that made the list. Some players are centered on a specific data type (e.g. Rasa and Hugging Face for natural language), while others focus on different areas, such as productizing AI (e.g. Scale, Tecton, and Weights and Biases) or serving as the “compute layer” for running AI workloads (e.g. Anyscale). ELT & orchestration makes data movement possible. It is the transportation layer that ensures data reaches its destination accurately and on time. This category developed from the traditional ETL vendors that were built on on-premises drag-and-drop interfaces. The newer class of players, by contrast, is mostly cloud-native (e.g. Fivetran and dbt), developer-friendly (e.g. Astronomer and Prefect), and able to manage more complex dependencies across different data environments. Data governance and security are becoming essential concerns as the data stack grows more complex and more stakeholders become involved. Governance tools are needed — especially in highly regulated industries — to protect data and preserve compliance throughout the data lifecycle (e.g. OneTrust and Collibra). This category is relatively new and typically serves large enterprise companies that operate under regulatory oversight. Customer data analytics has traditionally been controlled by marketing teams. However, because of its growing importance, data teams are now more engaged in integrating customer data with central data platforms. This category is centered on capturing customer data (e.g. Rudderstack and ActionIQ) or operationalizing that data to support front-line business use cases (e.g. Census and Hightouch). BI & notebooks represent the consumption layer of data. Even though it is a mature category, newer players such as Preset or Metabase are taking an open source-first approach and attracting technical data engineers as well as business intelligence teams. The rapidly changing nature of data needs also creates more demand for iterative and interactive notebooks (e.g. Hex) and automated insight generation (e.g. Sisu). Data observability takes inspiration from best practices in the software engineering stack. As the data stack becomes increasingly interdependent on up- and downstream tooling, and the accuracy of data has a wider impact, observability emerged as the newest category to provide monitoring and diagnostic capability across the data flow.
Although the main market tailwind behind adoption is the rising volume and use of data, the underlying drivers vary by category. For instance, the advances taking place in querying and processing are primarily driven by the separation of compute and storage, the shift to the cloud, and lower-cost computing power. Meanwhile, the adoption of operational tooling in data governance and data observability is driven largely by growing operational use cases and the complexity of data workflows.
Capital raised by categoryQuery and processing companies have raised the lion’s share of capital
The query and processing category accounts for only one-fifth of the companies in Data50, yet the amount of capital — almost 50% of all funding — invested in this category is staggering. Even though this data is shaped by Databricks’s recent $1.6B funding round, the category would still represent 37% of all funding — more than twice that of the next category — without it.

When the categories are viewed by company count, the distribution is more even. AI/ML is the largest category by number of companies, largely because the space is still evolving and calls for a new separate set of tools to train, measure, and productionize models. (For more on how this space is evolving, read Emerging Architectures for Modern Data Infrastructure.)

Data50 companies by geographyThe Data50 is clustered in the Bay Area
Of the 50 companies, 47 (94%) are headquartered in the United States and three are international. Most of the companies, 33, are based in the San Francisco Bay Area, while nine are along the I-95 corridor in Washington, D.C., Philadelphia, New York, and Boston. Two are based in Seattle, one is based in Cincinnati, and one is based in Atlanta.
This pattern is strongly shaped by the historical location of the large-scale data ecosystem (Oracle and Teradata were both founded in the Bay Area, for example). Still, we’re seeing more data companies emerge around the world (e.g. Firebolt and Matillion) as data engineering talent and demand for data tooling reach nearly every continent.

Data50 companies by founding yearAI/ML category drove spike of new data companies in 2019
Most of the Data50 companies were founded after 2014, with a high point around 2019, fueled by the surge in AI/ML tooling. In fact, many more data companies were founded after 2019, but because we’re focused on companies that have reached a certain scale, most newer companies don’t appear on this list yet.

Investment dollars are rising in every category
When we look at investment by category, the clearest trend is that AI/ML companies are drawing more investor attention than ever, with most of it concentrated at the early stage. The same is true for ELT and orchestration – largely propelled by mega rounds from Fivetran and dbt. Query and processing companies keep attracting large dollars, although the companies are generally later stage.

We strongly believe the next 10 years will be the decade of data, spanning infrastructure, applications, and everything in between. Because of that, we’ll keep seeing record-breaking growth, funding, and market capitalization, which we will track each year in this list. Congratulations to all the companies in the first Data50 class!