Earlier this year, Britton Stamper (cofounder of Push.ai and now at dbt Labs) published a LinkedIn article about how the definition of data has changed.
Not the dictionary definition, but how we think about data businesses can work with internally.
We’re already seeing the shift happen. Here’s what’s changing and how data teams need to respond.
More data is accessible than ever
Prior to the Modern Data Stack era beginning in the mid-2010s, teams often would only analyze and report on data from their application databases. Thanks to tools like Fivetran that made it easier to set up data pipelines, more teams brought in data from Salesforce, Zendesk, Stripe, and other popular sources.
Now with AI, Slack threads, call transcripts, and a bevy of other artifacts are now in play. Teams had this data available but didn’t have a way to reliably extract value from it.
AI changes that by making it much easier to work with natural language data.
Companies like Supersimple have even made it a big part of their value proposition. You can semantic search against both the data warehouse and Slack at the same time. If you’re trying to troubleshoot an unexpected value, you will probably check both these places anyway. Notably, this means that data teams need to be prepare for querying data that doesn’t live in a dbt model.
On the other hand, I also shipped a project at brightwheel where we scored sales call transcripts and added custom enrichment fields using Amazon Redshift and Amazon Bedrock. This hit parity with tools similar to Gong without the vendor lock-in, with more flexibility.
Querying data through AI agents requires context engineering
A key difference in querying data with AI agents vs. from a traditional BI tool or directly in the data warehouse is that agents need context.
That context can include a semantic or metrics layer, example queries and workflows, instructions on which tables to use and not use, style, and much more
If you’re using tools that have agent observability like Hex or Basedash, you should be reviewing conversation data. You’d be surprised just how bad an agent without context will mess up, while confidently presenting an answer to the user.
Armed with the right context, you can see results like the Disco team after our engagement. Agent usage in Hex skyrocketed as trust developed within the business.
Which team should own the AI context layer?
Data teams have enough on their plate already. It’s natural to have some hesitation towards taking on ownership of the context layer.
However, Britton said it best:
“But here’s what makes this existential: if data teams don’t claim this scope, someone else will. Engineering, product, ops. They’ll build the context layers, own the AI strategy, and the data team will stay behind answering “what were my top five products last month?” while the most important work in the organization happens without them.”
If data teams allow another team to own the context layer, they’ll find themselves 1) lacking the agency to update or contribute to the context layer 2) disconnected from how users interact with data.
There may need to be a partnership between data and a technical team like engineering or IT, but data should be leading the charge, not a bystander.
Conclusion
The context layer is the most important thing for data teams to own in the near future. There should be serious conversations about what roadmap projects and recurring work need reprioritizing to make room.
The job of the data team will look different next year than it did last year.
Need help making the transition. Let’s chat.