Wednesday, July 29, 2026

Is Databricks and Fabric Overtaking Snowflake?

Everywhere you look, companies are talking about Databricks or Fabric. Three years ago it was all Snowflake. So is Snowflake losing? Short answer from a working thread of engineers, consultants, and people inside the buildings: no — but the question itself is the wrong shape.

RAW DATASNOWFLAKEwarehouse · governed SQLDATABRICKSlakehouse · ML / governanceFABRICbundled · Azure onrampworkload fitEDW / BI · analysts, finance, reportingML pipelines · data science, AI
FIG. 01 — routing pattern reported across the threadmost orgs split by workload, not brand loyalty
A note on what this is. This synthesizes a discussion among data engineers, consultants, a self-described current Snowflake employee, and a Fabric user who ran a two-year enterprise pilot. It's lived, anecdotal, sometimes contradictory experience — not vendor benchmarks, audited market share, or verified financials. Specific numbers raised in conversation (stock moves, usage ratios) are presented as one person's claim, not confirmed fact.
01 — The direct answer

Is Snowflake losing the race?

Every practitioner in this thread converged on roughly the same answer, from different angles.

Verdict

Snowflake isn't losing — the race itself was never a single race. Databricks is growing faster and winning more greenfield deals in specific situations (ML-heavy teams, regulated industries needing BYOC). Snowflake still has the larger, more mature installed base as a governed SQL warehouse. Fabric is gaining share for reasons almost entirely unrelated to product merit. None of that is "overtaking" in the way the hype cycle implies.

What actually changed in three years isn't who's winning — it's who's talked about. Hype rotates on its own schedule, decoupled from revenue, retention, or platform quality. The clearest evidence for this in the thread: Oracle has a $350B market cap and Snowflake has roughly $95B, and nobody is out there hyping Oracle. Oracle is not losing. It just exited the "exciting growth story" phase of its life decades ago. Snowflake and Databricks are still in that phase — which is exactly why they generate this much conversation.

02 — Why it feels like Databricks is winning

Real reasons, not just hype

Origin advantage

Governance & ML

Databricks grew out of Spark and machine learning, where lineage and reproducibility were existential from day one. Unity Catalog is the product of that history — governance was foundational, not bolted on.

Architecture advantage

Data sovereignty

Classic Databricks architecture keeps the data plane — clusters and the data itself — inside the customer's own cloud account. For regulated industries, "our data never sits in a third party's storage" isn't a preference, it's a disqualifier for anything else.

Paradigm advantage

Lakehouse became dominant

The lakehouse paradigm — one copy of data, many engines — became the industry's reference architecture. Databricks was the original lakehouse company, which gives it a "we called it" credibility Snowflake has had to retrofit its way into.

One more data point from the thread, offered with a caveat: a dbt representative reportedly estimated dbt usage runs roughly 3x higher on Snowflake than on Databricks. The read on this wasn't "Databricks is worse" — it's a different paradigm. Snowflake users lean on dbt because Snowflake's SQL-warehouse users needed an external tool to get version control and transformation discipline. Databricks users more often build their own pipelines natively, so a dbt-shaped tool is optional rather than load-bearing. A migration story mentioned in the thread — Asana reportedly moving from Snowflake to Databricks — was flagged by the person raising it as an interesting but somewhat self-promotional case study, worth reading skeptically rather than as proof of a trend.

03 — Capability sentiment

What the thread said each platform is actually good at

Illustrative, not measured — a rough plot of where consensus landed across six axes practitioners kept returning to. 

SQL WarehouseML / AI ToolingGovernanceCost PredictabilityData SovereigntyBundling Pull
SNOWFLAKE
Warehouse-first, cost-predictable, weak on sovereignty
DATABRICKS
ML-first, strongest governance, runs in your cloud
FABRIC
Wins on distribution, not on any single axis
04 — How they got here

Convergence, not collision

Each platform's current edge traces back to its origin story — and the rivalry has visibly accelerated both roadmaps.

EARLY DAYS
Snowflake: the cloud data warehouse
Built for fast, governed SQL over structured data. No native git or CI/CD for years — dbt filled that gap. Analysts loved it; engineers had to route around it.
EARLY DAYS
Databricks: Spark, notebooks, ML
Grew out of data science. Git-friendly from the start, because reproducibility was existential to the job it was doing from day one.
THE PUSH
Both add what the other had
Snowflake ships Snowpark, Cortex, and closer git/CI-CD integration, and gets visibly closer to Databricks-style workflows. Databricks ships Databricks SQL and Unity Catalog to compete on warehousing and governance. Iceberg becomes the shared battleground for open, customer-owned storage.
MEANWHILE
Microsoft assembles Fabric
Synapse, Power BI, and Data Factory get folded into one brand. Old Synapse/ADF jobs can be "mounted" into Fabric while native functionality catches up — everything runs on ARM templates underneath, so Microsoft can keep swapping the UX layer for the next rebrand once parity is reached, and paying-support customers reportedly get an "army of engineers" to help migrate.
NOW
Same job, different starting points
Practitioners in this thread largely agree: in 2026, Snowflake and Databricks can both do most of the same things well. The differentiator has shifted from raw capability to architecture — who owns the compute, who owns the storage, and what that means for regulated data — and to which team's background (SQL analyst vs. ML engineer) each platform still naturally fits best.
05 — Where Fabric actually stands

Growing share, not growing on merit

Nobody in the thread — including an actual Fabric user who ran a two-year enterprise pilot — thinks Fabric is competing with Snowflake and Databricks on product quality.

Synapse + ADFlegacy jobs, already sold"mounted" intoFabric UX layerfunctionality catching upruns onARM templatesinfra underneath, brand-agnosticenablesNext "Fabric," 5 yrs outsame infra, new brand, new pitch

The pattern described: Fabric is genuinely useful for small teams or for extending what a single data analyst can do — a real, if narrow, use case. At enterprise scale, multiple people in the thread reported the opposite experience: a two-year enterprise pilot as a preferred partner concluded it "simply wasn't ready," with inconsistent access control between modules — visible evidence that different Microsoft teams built different parts of Fabric without lining up on fundamentals. One practitioner's summary: "give it a couple of years, the sales team would have sold it to enough people to make it worth using" — i.e., the product may eventually earn its share, but isn't earning it today.

Nobody in the thread has seen a company move from Snowflake to Fabric. Fabric's growth comes almost entirely from companies already locked into Azure licensing, where it shows up as "free" — a framing repeatedly called misleading, since the real cost shows up later as lock-in. It was described bluntly as "a tool that Finance makes you use, or a service provider that's only trained on it — it never wins on merit."

06 — Voices from the thread

What people actually said

"I think Snowflake have gotten a bit obsessed with beating Databricks, rather than just being a great product by itself."— practitioner, extensive user of both
"Their first all-hands, people asked about Databricks this-and-that, and the new CRO basically said: I don't give a **** about Databricks. I want us focusing on having a great platform."— self-identified current Snowflake employee
"Databricks is a Microsoft partner and native to its ecosystem, so it's an easy step up."— Fabric user, two-year enterprise pilot
"You can literally tell different teams have made different parts of it — none of them line up, especially access control."— practitioner, ran a Fabric preferred-partner pilot
"Snowflake runs its own storage and its own compute — that's not even an option for us from a security standpoint, regardless of compliance certifications."— practitioner in a regulated industry
"There's no buzz about Oracle and its market cap is $350B. Oracle is not losing."— consultant, platform advisor
"The split came down to Iceberg interoperability. Snowflake's Horizon catalog wrote to external Iceberg tables cleanly — Databricks' Unity Catalog couldn't, reliably, at the time."— consultant, hybrid-architecture case study
"Fabric is a tool that Finance makes you use, or you have a service provider that's only trained on it. It never wins on merit."— consultant, weekly platform advisory
07 — What most teams actually do

Split by workload, not by brand

Case note — consulting engagement

Snowflake for the warehouse, Databricks for ML pipelines, on the same client. Neither platform "lost" — the decision was made per workload.

Why Snowflake, here

Horizon catalog wrote to external Iceberg tables cleanly. Billing was predictable for this workload's shape — no cluster tuning required to avoid overpaying.

Why Databricks, here

ML pipelines needed the notebook/Spark-native environment. Unity Catalog's Iceberg interop was less reliable at the time, so it stayed out of that specific job.

A related bet raised in the thread: some see Databricks pulling ahead structurally because of Genie and an emerging ontology/semantic layer — the idea that once natural-language BI is good enough, companies will prefer one platform with a unified semantic layer over stitching meaning across multiple tools. That's a bet on convenience and unification beating best-of-breed — notably, the same structural bet Fabric is making through bundling, just backed by stronger underlying engineering, in this framing.

08 — Where Snowflake is showing friction

Real product complaints, not just narrative

Two concrete gripes came up that are worth separating from the culture/stock noise, because they're about the product itself.

Dynamic Table failures

Diagnostics, not just workarounds

Recent DT errors got a "split up the query" response rather than a root-cause explanation. Once you're manually decomposing a Dynamic Table's logic to work around engine limits, you've lost the point of using one over a stored procedure with explicit control flow.

No native modeling tool

ERDs still live outside Snowflake

No confirmed agreement with any third-party modeling vendor exists — but Snowflake's broader pattern is to stay narrowly focused on the core engine and leave adjacent categories (BI, orchestration, data modeling) to ecosystem partners. Defensible strategy; leaves a real gap as semantic layers become more central.

09 — On culture and stock price

A voting machine, not a weighing machine

The original question leaned on a single former employee's account: culture gone bad, can't retain good people, stock "keeps crashing." What came back from people closer to it was more specific and more mixed.

  • A quarterly review program was described as having become a disguised layoff mechanism, unevenly applied — including against people on leave or of a certain age — a real, specific complaint, not a vague vibe.
  • That program is reportedly no longer in use, following internal morale and culture problems it caused.
  • A current employee of seven years reports never having had a bad manager, framing "constant change" as a byproduct of scaling from roughly 1,100 to over 9,000 employees.
  • A new head of sales reportedly told an all-hands to stop obsessing over Databricks and focus on the product — read internally as an active course-correction.

Both things can be true in a company that size: a specific program was a real problem for a period, and the aggregate day-to-day experience is reported as good by longer-tenured staff. One data point doesn't settle it either way.

On the stock: one participant cited SNOW as up roughly 42% in 2025 and roughly 22% so far in 2026 — figures from the conversation, not independently verified here. Either way, short-term price moves mostly reflect growth-rate expectations against a stock originally priced for hyper-growth, not platform quality. The better question for anyone actually choosing a platform: is it stable, does it deliver value for the spend — not where the share price sat last quarter.
The rivalry is the R&D budget. Overtaking was never the right frame.

Five years of Snowflake and Databricks pushing each other produced two platforms that, per this thread, now do most of the same things well — and that pressure is a large part of why. Fabric wins procurement battles it hasn't earned on merit, and may keep doing so for years on distribution and bundling alone, the same way Databricks may be gaining ground with a semantic-unification bet that isn't purely a technical one either. None of that settles a leaderboard. It confirms the original question was under-specified: the honest answer is workload fit, data-sovereignty requirements, and team background — not who's "overtaking" whom.

Data Analytics advances in the age of LLMs

 

The rise of large language models is not only reshaping data engineering — it is fundamentally changing how we analyse data. For decades, analytics has been the domain of specialists writing SQL, building dashboards, and training models. Today, LLMs are turning that model on its head, enabling natural language interfaces, automated insight generation, and democratised analytics for everyone.

In this article, we explore the key advances that are powering the next generation of data analytics: from text‑to‑SQL and augmented analytics to automated data preparation and prescriptive recommendations. But we also take a hard look at the risks: over‑reliance on AI, data quality issues, and the organisational debt that can turn AI‑powered analytics into a costly failure.

✦ ✦ ✦

1. Natural language queries: the end of SQL as we know it?

One of the most visible impacts of LLMs on analytics is the ability to ask questions in plain English and get answers directly from data. Tools like ThoughtSpot, Power BI Copilot, and Snowflake Copilot are now using LLMs to translate natural language into SQL, MDX, or even Python — making analytics accessible to business users who never learned to code.

This is not just a convenience; it's a paradigm shift. The barrier to entry for data exploration has dropped dramatically. A product manager can now ask, “What were our top‑selling products last quarter in the EU?” and get a chart in seconds, without waiting for a data analyst.

Natural Language to SQL — querying with plain EnglishUser Query“Show me sales by region”LLMNLP → SQL(with schema awareness)Generated SQLSELECT region, SUM(sales) …Result / ChartInstant visualisationLLMs can generate SQL, but they also need context, schema, and validation to avoid errors.
Figure 1: LLMs bridge natural language and SQL, enabling self‑service analytics — but accuracy depends on data context and validation.

However, this power comes with risks. LLMs are notorious for hallucinating incorrect SQL, especially with complex joins or ambiguous schema. Data analysts still need to validate, optimise, and debug the generated queries. The skill of writing efficient SQL is not obsolete — it's evolving into a review and refinement role.

✦ ✦ ✦

2. Augmented analytics: AI‑powered insights

Beyond querying, LLMs are being used to automatically generate insights from data. This is often called augmented analytics — where AI suggests correlations, anomalies, trends, and even root causes without being asked. For example, an LLM might scan a sales dashboard and highlight that “Sales in the Midwest dropped 15% this month, driven by a decrease in repeat customers.”

This capability turns analytics from a pull (ask a question) to a push (get proactive alerts) model. It can save time, uncover blind spots, and help less technical users understand what's happening in their data.

Augmented Analytics — automated insight generationDataInsightEngine(LLM + statistical tests)Generated Insights• Sales dip in region X• Correlated with promo end🔁 Analyst validates, refines, and learns
Figure 2: Augmented analytics uses LLMs to proactively surface insights, but human validation remains essential.

The challenge is that these insights are only as good as the data and the model. Spurious correlations and misleading narratives can easily creep in. Therefore, organisations must still invest in data literacy and critical thinking — skills that AI cannot replace.

3. Automated data preparation and cleaning

Data scientists and analysts spend up to 80% of their time on data preparation — cleaning, transforming, and joining datasets. LLMs are now being used to automate many of these tasks. For instance, an LLM can suggest data type conversions, missing value imputation, or even generate Python/Pandas code to reshape a messy CSV.

Tools like Pandas AI, Trifacta, and Alteryx are integrating LLMs to recommend transformations based on natural language descriptions. This significantly speeds up the “data wrangling” phase, allowing analysts to focus on higher‑value work.

However, this automation is not foolproof. LLMs can misinterpret column semantics or apply transformations that introduce bias or data leakage. Again, the human in the loop must review, test, and document every step to maintain data integrity.

Automated CleaningMessy CSVLLM clean(impute, format)Clean dataset⚡ 80% time saved
🧹 Automated data prep
Feature EngineeringRaw featuresLLM suggest(new features)Enhanced set💡 ML‑ready
⚙️ Feature suggestions
✦ ✦ ✦

4. Predictive and prescriptive analytics with LLMs

LLMs are not just for querying and cleaning; they are also being used to generate predictive models and prescriptive recommendations. For example, an LLM can take a dataset and suggest which machine learning model to use, or even write the training code. More advanced systems can explain model predictions in plain English, making AI more interpretable.

Prescriptive analytics goes a step further: the LLM can recommend actions based on predictions. For instance, if the model predicts a high churn risk, the LLM might suggest “Send a 10% discount coupon to these customers” and even draft the email copy.

This blurs the line between analytics and decision‑making. It also raises ethical concerns — who is accountable for automated decisions? Data analysts must become AI ethicists and risk managers as they deploy these systems.

“Prescriptive analytics powered by LLMs can drive business value, but they also require robust governance, transparency, and human oversight.”
✦ ✦ ✦

5. The double‑edged sword: LLMs and analytics skills

Just as in data engineering, the rise of AI in analytics carries the risk of skill erosion. When a chatbot can write SQL, generate a chart, or even build a model, what happens to the analyst's ability to think critically, understand data lineage, and question assumptions?

There is a growing concern that over‑reliance on LLMs will create a generation of analysts who can prompt but cannot debug. For instance, asking an LLM to “find the root cause of sales decline” might produce a plausible explanation, but without domain knowledge and statistical rigor, that explanation could be completely wrong.

Moreover, the search and synthesis skills that analysts traditionally develop by reading forums like Stack Overflow or academic papers are being bypassed. The act of struggling with a problem often leads to deeper understanding — and that struggle is being outsourced to a machine.

Analyst + LLM — collaboration or dependency?AnalystLLM(assistant)Insight / SQL(needs validation)⚠️ The analyst must retain
critical thinking, domain knowledge, and statistical literacy — AI is a tool, not a replacement.
Figure 5: The partnership between analyst and LLM is powerful, but over‑dependency can erode core skills.

The solution is to use LLMs as junior assistants — they generate drafts, suggest alternatives, and automate routine tasks. But the final validation, interpretation, and storytelling must remain human. This requires a deliberate effort to maintain analytical rigor and continue investing in training.

✦ ✦ ✦

6. The organisational and cultural reality of AI‑driven analytics

“If your company has never cared about clean data, then AI integration at scale will not be feasible without starting from the bottom and working up. If your company doesn’t have an engineering culture and you let people enter whatever they want into your CMS/financial system/whatever, then you’re not going to get what you expect out of AI.”

“The AI kool‑aid crazy has been getting smothered by research pointing out that AI shifts the workload in many cases, blunts talent in other cases — the outages, the dropped databases, and the financial losses. … Take something like generating takeoffs from engineering drawings and building price sheets. That’s not exactly pushing the limits of modern AI. But if your inputs are inconsistent, your pricing data is all over the place, and nobody owns the implementation, it’s almost guaranteed to fail.”

These words, echoed across many industries, highlight a painful truth: AI cannot fix broken data culture. Many organisations rush into AI‑powered analytics without addressing the fundamental hygiene of their data. Inconsistent formats, missing values, duplicate records, and poor documentation become amplified when fed into an LLM — leading to garbage‑in, garbage‑out at scale.

Moreover, the cultural shift required is often underestimated. Analysts and business users need to learn how to prompt effectively, interpret outputs skeptically, and take ownership of the final product. Without this, the organisation becomes dependent on consultants who build black‑box solutions that fail after they leave.

The financial risks are real. Millions of dollars have been wasted on AI projects that were poorly scoped, poorly governed, and poorly executed. The privacy concerns around feeding sensitive data to cloud‑based LLMs add another layer of risk. And the recent backlash against massive data center investments suggests that the ROI of AI is far from guaranteed.

For analytics leaders, the path forward is clear: invest in data governance, data literacy, and a culture of quality before deploying AI at scale. Treat AI as an accelerator, not a silver bullet. And always keep a human in the loop — because the most powerful insights are those that combine machine intelligence with human wisdom.

✦ ✦ ✦

7. The road ahead

As LLMs continue to evolve, we can expect even more advanced analytics capabilities:

  • Multi‑modal analytics: Combining text, images, and tables in a single query.
  • Autonomous agents: LLMs that can not only answer questions but also take actions (e.g., send emails, update dashboards).
  • Continuous learning: Systems that improve over time by incorporating feedback from users.
  • Explainable AI: LLMs that can provide detailed reasoning behind their suggestions, building trust.

But with these advances comes a greater need for responsible AI — ensuring fairness, transparency, and accountability. Analytics professionals will need to upskill in AI ethics, model governance, and change management.

The future of analytics: human + AIHuman Analystcuriosity, contextLLMscale, speed, automationBetterDecisionsThe partnership between analysts and AI will define the next decade of analytics.
Figure 6: The future is not AI replacing analysts, but AI empowering them to be more effective.
✦ ✦ ✦

The age of LLMs is ushering in a golden era for data analytics. Natural language interfaces, automated insights, and intelligent data preparation are making analytics more accessible, faster, and more powerful than ever before. But this power must be wielded with responsibility, critical thinking, and a strong foundation in data quality and governance.

As analysts, we must evolve from being query writers to insight curators and AI supervisors. We must embrace the tools, but never lose sight of the art and science of asking the right questions, interpreting results, and communicating narratives that drive action.

The organisations that will succeed are those that invest in people, processes, and culture — not just technology. They will treat AI as a force multiplier for human intelligence, not a replacement. And they will remember that the most important analytics skill is still curiosity.

Further reading:
How to build a natural language query engine with LLMs
The ROI of augmented analytics: case studies
Data governance for AI‑powered analytics

#DataAnalytics #LLMs #NLPtoSQL #AugmentedAnalytics #AIethics

Tuesday, July 28, 2026

Data Engineering advances in the age of LLMs

The explosion of large language models has done more than just push the boundaries of natural language understanding — it has fundamentally altered the landscape of data engineering. The pipelines that once fed simple dashboards and ML models now must serve real‑time, high‑dimensional, and semantically rich data to LLMs at scale.

In this article, we explore the key advances that are defining the new data stack: from vector databases and embedding pipelines to streaming RAG and data freshness challenges. But we also take a sober look at the organisational, cultural, and financial realities that can turn AI adoption into a costly failure. If you're a data engineer, architect, or AI practitioner, this is your guide to both the frontier and the pitfalls.

✦ ✦ ✦

1. The rise of vector databases

Traditional OLTP and OLAP systems were built for structured, scalar data. LLMs, however, consume embeddings — dense vector representations of text, images, or audio. This shift has given birth to a new category of databases optimized for similarity search and approximate nearest neighbor (ANN) queries.

Vector databases like Pinecone, Weaviate, Milvus, and pgvector have become first‑class citizens in the data engineer's toolbox. They allow you to store billions of vectors and retrieve the most relevant contexts for an LLM in milliseconds — a critical requirement for RAG (Retrieval‑Augmented Generation) systems.

Vector Database — embedding storage & retrievalDocs / ChunksEmbedding Model(e.g., text‑embedding‑3)Vector DBindex + metadataANN Querytop‑k similarv₁v₂v₃v₄v₅v₆v₇★ nearestHNSW · IVF · PQ
Figure 1: A vector database stores embeddings and answers ANN queries to retrieve the most relevant contexts for an LLM.

But vector databases are not just a new storage layer — they demand a rethinking of indexing strategies, sharding, and replication. Data engineers now need to think about embedding freshness, dimensionality trade‑offs, and the cost of approximate search. This is a far cry from the B‑tree and hash indexes of yesteryear.

2. RAG pipelines: the new ETL

Retrieval‑Augmented Generation (RAG) has become the de facto pattern for grounding LLMs in proprietary or up‑to‑date data. A RAG pipeline is essentially a data engineering workflow that ingests, chunks, embeds, and stores documents — then orchestrates retrieval and generation at query time.

The diagram below shows a typical RAG pipeline. Notice how it blends batch (indexing) and streaming (query) concerns — a hybrid model that many data teams are now adopting.

RAG Pipeline — indexing + query flowINDEXING (batch)DocsChunking+ cleaningEmbedmodelVector DB(store)QUERY (real‑time)UserqueryEmbedqueryANN Searchtop‑k contextsLLM + RAGgenerateOutput: grounded, context‑aware response hybrid batch + stream
Figure 2: RAG pipelines blend batch indexing (top) with real‑time query processing (bottom) — a new data engineering paradigm.

From a data engineering perspective, RAG introduces several new challenges:

  • Chunking strategies: How do you split documents to preserve semantic coherence while respecting token limits?
  • Metadata filtering: Combining vector similarity with structured filters (e.g., date, source, department).
  • Freshness: How do you keep the vector index in sync with source data changes?
  • Evaluation: Measuring retrieval quality with metrics like hit rate, MRR, and NDCG.
✦ ✦ ✦

3. Streaming & data freshness

LLMs are trained on static snapshots, but the world changes fast. Data engineering in the LLM age is increasingly about freshness — how do you bring the latest information into the model's context window? This has led to a surge in interest around streaming RAG and incremental indexing.

Modern data stacks are adopting change data capture (CDC) from operational databases, feeding into streaming platforms like Kafka or Redpanda, and then into vector DBs with upsert capabilities. This allows RAG systems to reflect updates within seconds rather than hours.

Streaming RAG — freshness at scaleCDC Source(Postgres, MySQL)Kafka(change stream)Embed + Index(incremental)Vector DB(always fresh)⏱️Freshness: sub‑second latency from source change to queryable embedding CDC → stream → embed → upsert → ready
Figure 3: Streaming RAG pipelines enable near‑real‑time freshness, ensuring LLMs always have the latest context.

This shift requires data engineers to become proficient with streaming stateful processing, watermarking, and exactly‑once semantics — skills that were once the domain of real‑time analytics teams but are now central to AI infrastructure.

4. Data quality for LLMs

"Garbage in, garbage out" has never been more true. LLMs are sensitive to the quality, diversity, and representativeness of the data they ingest. Data engineers now play a critical role in ensuring that the data feeding RAG systems and fine‑tuning pipelines is clean, unbiased, and fresh.

New tools and practices are emerging: data profiling for embeddings, outlier detection in vector space, and automated data drift monitoring. The goal is to build trust in the data that powers LLM applications.

Embedding Drift▲ drift detected
🔍 Drift monitoring
Data LineageSourceEmbedmodelVectorDBLLMRAG
📊 Lineage tracking

Data engineers are also adopting data contracts and schema registries for embedding pipelines, ensuring that changes to the embedding model or chunking strategy are backward‑compatible and well‑documented.

“The quality of an LLM's output is bounded by the quality of the data it retrieves. Data engineering is now the quality gate for AI.”

5. The double‑edged sword: LLMs and data engineering skills

While LLMs are revolutionizing data pipelines, they also bring a subtle risk: the erosion of fundamental engineering skills. Data engineering is a deeply technical field that requires understanding of distributed systems, data modeling, performance tuning, and fault‑tolerance. Yet, the rise of AI coding assistants has led many to lean on chatbots for everything from writing Spark jobs to debugging SQL.

Historically, tools like MATLAB have been used for quick data analysis and cleaning—often in research or scientific settings. But that is not data engineering. Data engineering is about building scalable, reliable, and maintainable data infrastructure — not just crunching numbers in a REPL. Python has become the lingua franca of the field because of its rich ecosystem (Spark, Airflow, dbt, etc.) and its seamless integration with ML frameworks.

The current “chatbot hype” is worrying some practitioners. It’s becoming common to ask a chatbot to write code for everything — even for generating MATLAB scripts that scrape Stack Overflow for answers. While this can boost short‑term productivity, it risks creating a generation of engineers who cannot debug, optimize, or design systems from first principles. The skill of searching, reading, and synthesising answers from forums like Stack Overflow is itself a valuable learning process that is being bypassed.

Consider this example: a data engineer might ask a chatbot to write MATLAB code that searches Stack Overflow for solutions to a specific error. The chatbot returns a snippet that uses the MATLAB webread function to query the Stack Exchange API. While that snippet may work, the engineer never learns about API rate limits, proper error handling, or how to parse JSON responses — all essential skills.

Chatbot‑generated code & the risk of skill erosionEngineer“Write MATLAB code to search SO”LLM(chatbot)MATLAB code(webread, API call)⚠️ Short‑cut productivity vs. long‑term mastery —use chatbots as assistants, not replacements
Figure 5: While chatbots can generate code (e.g., MATLAB for Stack Overflow queries), over‑reliance can weaken core debugging and design skills.

The key is to treat LLMs as force multipliers, not substitutes. A skilled data engineer uses a chatbot to generate boilerplate, explore alternative syntax, or quickly prototype — but always reviews, understands, and tests the code. The same applies to MATLAB, Python, or any other language. The real value lies in the architectural thinking, trade‑off analysis, and operational knowledge that no chatbot can replicate.

As we embrace LLMs in our data stack, we must also invest in continuous learning and hands‑on practice. The hype will fade, but the fundamentals of data engineering — reliability, scalability, and correctness — will remain.

✦ ✦ ✦

6. The organisational and cultural reality of AI adoption

“That is a risky position to take in the current environment. If there is a tool that can make you more efficient at your job and you avoid using it because you prefer to do it the manual way, you risk becoming ineffective compared to your peers.

There are real limitations and risks with using AI that you should understand and account for, but having it save you time by doing a lot of the research for you is one area that it generally outperforms us humans.

As someone who enjoys learning and the academic side of things, this can be a difficult part of the job to delegate to a machine, but if it makes you more effective and valuable to your employer - it is a trade off worth considering.”

These words, shared by a seasoned practitioner, capture the tension many of us feel. On one hand, AI tools are undeniably powerful accelerators. On the other, they are not a silver bullet — and when deployed without the right cultural and organisational groundwork, they can become expensive, risky distractions.

Manufacturing companies, for example, are constantly going over budget trying to implement AI tools in an environment that is not suitable for it and with employees who do not have the proper cultural training to keep those tools accurate. I’m glad you’re seeing gains with your projects but that’s not really relevant when you’re thinking on org‑wide scales.

The harsh reality is that if your company has never cared about clean data, then AI integration at scale will not be feasible without starting from the bottom and working up (which they always want to skip because they want AI yesterday, not months or years from now). If your company doesn’t have an engineering culture and you let people enter whatever they want into your CMS/financial system/job management tool/whatever, then you’re not going to get what you expect out of AI.

The AI kool‑aid crazy has been getting smothered by research pointing out that AI shifts the workload in many cases, blunts talent in other cases — the outages that have been blamed on AI, the dropped databases, and the financial losses. There are privacy concerns with giving your trade secrets in natural language to cloud companies. Every podcaster and their momma is talking about how the data center investments don’t make sense, how AI is facilitating a modern day “wealth transfer,” infranoise… Even Uber has been starting to pull the rug out a bit.

Take something like generating takeoffs from engineering drawings and building price sheets. That’s not exactly pushing the limits of modern AI. But if your inputs are inconsistent, your pricing data is all over the place, and nobody owns the implementation, it’s almost guaranteed to fail.

I think a lot of executives underestimate how much organisational debt they’re carrying into these projects as well as underestimate how much of a culture change is required to really succeed with them. Then consultants spend months trying to paper over those problems until leadership decides to pull the plug, millions of dollars later.

“The AI hype will fade, but the underlying data mess will remain. Clean data, clear ownership, and a culture of quality are prerequisites — not nice‑to‑haves.”

This does not mean we should abandon AI. Rather, it means we must approach it with eyes wide open. The most successful organisations will be those that invest first in data governance, engineering excellence, and continuous training — and then layer AI on top. They will treat AI as a powerful assistant, not a magic wand, and they will measure success not by the number of chatbots deployed, but by the reliability, trustworthiness, and business value of the outcomes.

✦ ✦ ✦

7. The road ahead

As LLMs grow more capable, the demands on data infrastructure will only intensify. We are already seeing the emergence of data‑centric AI platforms that unify data ingestion, transformation, embedding, and retrieval into a single cohesive stack. Data engineers are evolving into AI data engineers, with a deep understanding of both data systems and machine learning.

Key trends to watch:

  • Multi‑modal data: Images, audio, and video embeddings will become first‑class citizens.
  • Federated RAG: Querying across multiple vector DBs and data silos.
  • Self‑optimizing pipelines: Using LLMs themselves to tune chunking, embedding, and retrieval parameters.
  • Data observability for LLMs: Monitoring not just data quality but also retrieval effectiveness and generation quality.
The new data engineering stack for LLMsIngestionEmbedVector DBRAGLLMinferenceThe stack is evolving from batch ETL to real‑time, embedding‑first pipelines.
Figure 6: The modern data engineering stack is embedding‑first, real‑time, and RAG‑aware.
✦ ✦ ✦

The age of LLMs is not just an AI revolution — it's a data engineering revolution. The tools and practices that served us well for a decade are being rethought, rebuilt, and reimagined. As data engineers, we have the exciting challenge of building the data infrastructure that will power the next generation of intelligent applications.

Yet we must not forget the craftsmanship that underpins our field. LLMs are powerful allies, but they are not a substitute for systematic thinking, deep debugging, and architectural wisdom. Use them wisely, keep learning, and always question the output — because the data that flows through your pipelines ultimately shapes the decisions that matter. And above all, build on a foundation of clean data, strong governance, and a culture that values quality over speed. That is the only way to turn AI hype into lasting value.