How to invest in Databricks - the public investors are real but the stakes are too small to matter
Databricks raised $5B at a $190B valuation in Aug 2026 and is the only profitable name in the AI IPO pipeline. NVDA, MSFT, GOOGL and AMZN all hold stakes, but each is immaterial. The real listed play is the pure-play comparable, SNOW.
The standard "how to invest in Databricks" answer is "buy Nvidia, it holds a stake." That is technically true and practically useless: Nvidia's Databricks position is a rounding error against a multi-trillion-dollar market cap, so it moves NVDA not at all. Databricks sits in the awkward middle of this series. Unlike Stripe, it does have listed shareholders. Unlike OpenAI, none of those stakes is large enough to be a real proxy.
Databricks raised $5B at a $190B valuation in August 2026, and CEO Ali Ghodsi has said an IPO is more likely 2027 than 2026. This piece walks through what Databricks does, the valuation arc, the four public investors and why each is immaterial, and the one genuinely useful listed expression: the pure-play comparable, Snowflake.
The TL;DR. Databricks has real public investors (NVDA, MSFT, GOOGL, AMZN) but every stake is strategically small, so none is a needle-moving proxy the way MSFT is for OpenAI. The cleanest listed way to express the Databricks thesis is $SNOW, its closest public pure-play competitor, which is a peer and not a stakeholder, so it tracks the data-and-AI platform theme rather than Databricks' own mark. ARK Venture Fund holds a small direct sliver; accredited investors get secondaries. The one thing that would change all of this is a 2027 S-1, and Databricks is the rare AI-pipeline name that is already profitable.
What Databricks does
Databricks is the data-and-AI platform: the "lakehouse" that merges a data warehouse and a data lake into one system, plus Mosaic AI for building and serving models on top of a company's own data. Where Snowflake started as the cloud data warehouse and moved toward AI, Databricks started from Spark and machine learning and moved toward the warehouse. The two now compete head-on for the enterprise data-and-AI budget.
The distinctive fact for an investor is that Databricks is profitable, or close to it on the metrics that matter, which is rare in the AI-infrastructure cohort. Most of the 2026 AI IPO pipeline is deeply loss-making (see the OpenAI $14B forecast loss); Databricks is frequently described as the only profitable name in that pipeline. That is the crux of its eventual IPO pitch.
The valuation arc
| Round | Date | Valuation | Read |
|---|---|---|---|
| Series I | 2023-09 | ~$43B | NVDA takes its initial stake |
| Series J | 2024-12 | ~$62B | The AI-platform re-rate |
| Series K | 2025-08 | >$100B | First triple-digit mark |
| Strategic round | 2026-08 | $190B | $5B raise, Coatue / Blackstone / MGX / T. Rowe led |
Two observations. First, the markup is steep but tamer than the frontier labs: roughly 4.4x in three years, against OpenAI's and Anthropic's near-vertical arcs. Second, the recent rounds are increasingly financed by crossover and PE money (Coatue, Blackstone, T. Rowe, Point72, TPG) rather than pure VC, which is the funding pattern of a company being groomed for a public listing rather than one avoiding it. That is consistent with Ghodsi's 2027 signal.
The four public investors, and why none is a proxy
Databricks did what Stripe never did: it took strategic checks from listed companies. NVDA, MSFT, GOOGL and AMZN have all invested, NVDA since the $43B round in 2023. The problem for a proxy investor is proportion. Each of these is a multi-hundred-billion to multi-trillion-dollar company, and a strategic Databricks stake, even one that has 4x'd, is immaterial against that base. Owning $NVDA, $MSFT, $GOOGL or $AMZN for the Databricks exposure is like buying an index for one basis point of it: you get the exposure, but it will never move your position. This is the opposite of the Microsoft-OpenAI case, where the stake is ~27% and materially moves the stock.
So the honest ranking of Databricks exposure is not "which investor do I buy." It is "which listed business rises and falls with the same thesis."
The exposure map
1. Snowflake (SNOW) - the pure-play comparable
$SNOW is the closest listed expression of the Databricks thesis, not because it holds Databricks (it does not, they are rivals) but because it is the public company whose revenue rises and falls with the exact same enterprise data-and-AI budget Databricks competes for. If the thesis is "the data layer of the AI stack keeps compounding," SNOW is the way to own it on a public exchange today. The caveat is that it is a competitor, so a Databricks win can be a Snowflake loss and vice versa; it tracks the theme, not the private mark. Track SNOW live: /stocks/snow.
2. ARK Venture Fund (ARKVX) - the small direct sliver
$ARKVX holds a direct Databricks position among its private names, with no accreditation requirement, alongside its SpaceX, OpenAI, Anthropic and Stripe weights. Databricks is a minority of the fund, so this is a thin, diversified slice, with the interval-fund liquidity and NAV-lag caveats covered in the Anthropic piece.
3. The strategic investors - exposure without impact
NVDA, MSFT, GOOGL, AMZN. Real stakes, immaterial size, as covered above. Own them for their own theses (compute, cloud, ads), and treat the Databricks position as a free option that will never be large enough to notice.
4. Secondary markets (accredited investors only)
Forge, Hiive and EquityZen list Databricks shares (typically as SPV interests). Same mechanics as the rest of this series: accreditation required, $25-100K+ minimums, 3-5% fees, the last round as the reference mark. The only way to be long Databricks specifically before an IPO.
Where it sits in the AI stack
Databricks and Snowflake are the data layer of the AI-software stack: below the model labs (OpenAI, Anthropic) and above the raw compute (NVDA silicon, the hyperscaler datacenters). The structural bet is that whoever owns the enterprise's data owns the surface where AI actually gets deployed inside a company, because a model is only as useful as the proprietary data it can reach. That is why a data platform commands a $190B private mark in an AI cycle: it is the connective tissue between the model and the enterprise. It is also why the SNOW-versus-Databricks contest is worth watching as a two-horse race for that layer.
What to watch
- The 2027 S-1. The single catalyst. Ghodsi has guided to 2027; a filing turns the whole exposure map from proxy-and-wait into a buyable ticker, and Databricks' profitability makes it a cleaner debut than most of the pipeline.
- SNOW results as the read-through. Snowflake's enterprise consumption growth is the closest public tell on whether the data-and-AI budget Databricks depends on is expanding or contracting.
- The next private mark. Databricks reprices roughly annually; the direction of the next round is the read on private-market appetite for the name.
- Mosaic AI adoption. Whether Databricks' model-building layer wins share against standalone tooling is the tell on whether it is a data company bolting on AI or an AI company that happens to own the data.
- The crossover-investor mix. More PE and crossover money (Blackstone, T. Rowe, TPG) in each round is the pattern of a pre-IPO grooming. Bubble shifts and rule-based alerts on SNOW and the strategic investors are part of /pro.
Live data on the listed expressions: /stocks/snow · /stocks/nvda · /stocks/msft · /stocks/googl · /stocks/amzn - price, ETF holdings, bubble correlation, bot positions.
Bubble context: /bubbles/ai-software - the data-and-AI-platform cluster Databricks belongs to and how it's moving.
QuantAbundancia is educational research. Nothing here is investment advice. See /disclosures.
Related bubbles
Related research
Go deeper
Get the daily digest.
One email a day · alerts + bubble shifts + new research. Free during beta.
No spam. One email per day max. Pro adds Telegram trade alerts and higher AI-assistant limits.