Semantic Layer
AIData

Your semantic layer doesn't have to take a quarter

Michał PuchałaVP of DataJuly 19, 20267 min read

Skip the quarter-long build. Stand up a federated semantic layer in days, learn what questions matter, then industrialize only what earns it.

The standard advice for getting a fund's data "AI-ready" reads like a construction project. Inventory the sources. Stand up a warehouse. Build the pipelines. Model the semantic layer by hand, metric by metric. Govern it. Then, a quarter or two later, point an agent at it. The advice is not wrong. It is just slow in a way that quietly kills momentum, because the people who asked for answers have moved on by the time the platform is ready to give them.

There is a faster path to the same destination, and it changes the order of operations. You can stand up a working semantic layer across your existing sources in days, use it to learn which questions actually matter, and then industrialize the parts that earn it into a warehouse. The slow, careful build still happens. It just happens second, informed by evidence, instead of first, on a guess.

Disclosure up front: Vecten is an official partner of both Snowflake and TextQL, both of which appear in this post. Nobody is paying for it. We are describing the sequence we actually run for clients, and we would tell you if the tools didn't earn their place.

The problem with building the warehouse first

A data warehouse is an investment, and like any investment it pays off only when the thing you built gets used repeatedly. That is a great fit for the metrics a fund queries every week: portfolio ARR, ownership, reserves, the standing board pack. Consolidate those into Snowflake, govern them, serve them through Cortex Analyst, and you have something durable.

But a large share of the questions at a fund are not like that. They show up mid-diligence, they span four systems including two you will never justify ingesting, they get answered once, and they shape a decision. "We can answer that in three weeks, once we've built the connectors" is, for those questions, the same as no answer. Building the warehouse first optimizes for the repeated questions and starves the one-off ones — which is backwards, because the one-off questions are often the ones with a deal attached.

Integrate across sources without moving the data

This is the gap federated tools fill, and it is why TextQL sits in our stack. Its premise is no data migration: connect to the warehouse and the operational systems where they already live — CRM, billing, a PitchBook export, the spreadsheets the platform team maintains — and define a semantic layer that spans all of them. The ontology does the unifying work that pipelines would otherwise do, at a fraction of the upfront cost and time. (Their own account of why they abandoned direct text-to-SQL for an ontology architecture is worth reading: Why we use an ontology.)

The part that genuinely compresses the timeline is that you are not modeling that layer by hand from a blank page. You point the tooling at the connected sources and let it propose the objects, the attributes, and the joins — Company, Fund, Deal, Person; ARR, ownership, check size; which Deal belongs to which Company. AI does the first pass across many sources at once, and your data team's job shifts from authoring the model to reviewing and correcting it: fixing the entity resolution it got wrong, settling the metric definitions the investment committee will actually stand behind, overriding the field it guessed from when the real number lives in billing.

That is the right division of labor. The machine is good at the breadth — reading dozens of schemas and proposing a coherent first draft fast. The judgment calls — what "ARR" means for a healthcare company, which of three records is the real one — are yours, and they are where the durable value sits. You get to a usable semantic layer in days, and you spend your scarce expert time on the decisions that matter rather than on plumbing.

A generated ontology still has to survive next quarter

Generating that first draft is quickly becoming table stakes. Much of the category, warehouse vendors included, is converging on AI-assisted ontology generation — and far less attention is being paid to the harder half of the problem: what happens after week one. Schemas drift. A new billing system arrives. A restructuring changes what "ownership" means for two positions. An ontology that lives in a UI as a pile of settings quietly rots, and six months later nobody can say why a metric is defined the way it is or who changed it.

The approach we find credible treats the ontology as files and code. Definitions live in version control, changes go through review, and the model can be diffed and tested like any other engineering artifact. It is the same move dbt made for analytics engineering — bringing version control, testing, and code review to data transformations — applied one layer up, to business context. That is what makes maintenance tractable: when the ontology is code, keeping it correct is an engineering discipline instead of a standing meeting. It is also, we suspect, where the real moat in this category will settle — not in who generates an ontology fastest, but in whose ontology is still trustworthy two years in.

Then let the analyst agent work

Once that layer exists, an agent has something to reason against. TextQL's agent, Ana, plans, queries, runs Python in a sandbox, and shows its work — but the reason its answers hold up is not the agent, it is the layer underneath. Take any capable analyst agent — Ana included — and point it directly at raw schemas: it will write a syntactically perfect query against a semantically wrong understanding of your business, then hand you the result with full confidence. Point the very same agent at a governed ontology and you get answers an investment team can sign off on. The interface is identical; the trustworthiness comes from the model beneath it.

This is also the cleanest way to learn what to build next. Every question the team asks against the federated layer is a signal. The ones that get asked once and never again were never worth a pipeline. The ones that come back week after week are telling you exactly where to invest.

Industrialize what earns it, in Snowflake

We are a Snowflake partner, so the obvious question deserves a straight answer: doesn't Cortex already do this? For data that lives in Snowflake, largely yes — Cortex Analyst on a well-built semantic model is a strong product, and for a metric the whole team queries weekly, consolidating it into the warehouse and serving it through Cortex is the right end state. We build exactly that, and we recommend it.

The point is that consolidation is the destination, not the starting line. The questions that proved themselves during exploration earn their pipelines and move into Snowflake, where they run with full governance, lineage, and performance. Crucially, the semantic definitions you settled during the fast federated phase travel with them — the work of deciding what "ARR" means is done once and reused, not redone. One environment is where questions get tested cheaply; the other is where the winners get industrialized.

So the two tools are a sequence, not a rivalry. TextQL is how you integrate fast and find the signal. Snowflake is where the signal becomes infrastructure. Running them in that order is how you get a fund's data AI-ready in weeks instead of quarters, without building a single pipeline you'll later regret.

Where Vecten fits

Most of a Vecten Core engagement is the connective work this post describes: standing up the federated layer, supervising the AI-assisted modeling so the ontology is actually right, resolving entities across CRM, warehouse, and third-party data, locking the metric definitions and keeping them governed as the systems underneath change, and deciding which questions graduate into Snowflake. Whether a given answer is ultimately served through TextQL, through Cortex, or through whichever agent is best in two years is a deployment decision, and a reversible one. The model of the business is the part that compounds — and now you can get the first version of it in days.

THE NEXT LEVEL

Let's talk about what
you're building

An AI-native partner that's already done to itself what it now does for its clients.

ENGINEERING
15 years depth
CLIENT AUM
$1.2 trillion+
NPS SCORE
80+
PARTNERSHIP
AI-native partner