38 episodes
- AI can now write SQL, generate Python, review pull requests, and build entire applications with a few well-crafted prompts. Yet despite all of that progress, data engineering still feels oddly resistant to the full AI revolution. Pipelines break. Data models drift. Production data needs validation. Someone still has to close the loop.
In this episode of the Data Engineering Central Podcast, I sat down with Hugo Lu, founder and CEO of Orchestra, to talk about what “agentic data engineering” actually means beyond the buzzwords. Rather than another conversation about AI replacing engineers, we dug into the infrastructure that’s still missing before autonomous data platforms become reality.
* Hugo shares his unlikely path into data engineering, from investment banking to helping build data systems at Juul, before eventually founding Orchestra.
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
What started as an effort to simplify orchestration has evolved into a broader vision where AI agents don’t just generate code, but can safely execute work, observe the results, validate changes, and iteratively improve pipelines inside secure environments.
* That ability to observe outcomes, what many are calling “closing the loop,” may be the missing ingredient preventing today’s coding agents from becoming truly autonomous.
We also explore why data engineering has not experienced the same AI disruption as traditional software engineering. While AI can produce application code remarkably well, production data systems introduce a completely different set of problems. Branching production data, validating schema changes, testing transformations against realistic datasets, and understanding business semantics all remain difficult challenges that require much more than simply generating code.
The conversation naturally turns toward the future of the profession itself. We discuss whether junior engineers are losing the traditional apprenticeship path, how senior engineers are shifting from writing code to reviewing AI-generated work, and why decades of experience debugging production systems may actually become even more valuable in an AI-first world. Rather than eliminating engineering expertise, AI may simply be changing where that expertise is applied.
Data Engineering Central is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
Finally, we dive into the rapidly changing data infrastructure landscape. From DuckDB and Polars to serverless compute, Iceberg, semantic layers, AI-native orchestration, and the growing concern over rising LLM token costs, we discuss where the industry appears to be heading and which trends are likely to stick long after today’s hype cycle fades.
If you’ve been wondering what comes after Orchestration, how AI agents will actually manage production data pipelines, or whether data engineering itself is about to undergo its biggest transformation in a decade, I think you’ll enjoy this conversation.
In this episode we discuss
* What “agentic data engineering” actually means.
* Why AI still struggles to fully automate data engineering.
* The importance of closing the loop with production feedback.
* Why orchestration may become the operating system for AI agents.
* The changing role of data engineers in an AI-first world.
* Why junior engineers face a very different career path than previous generations.
* DuckDB, Polars, serverless data platforms, and where modern infrastructure is heading.
* Whether today’s dependence on proprietary LLMs will create tomorrow’s vendor lock-in.
* How Orchestra is building infrastructure for AI-native data platforms.
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit dataengineeringcentral.substack.com/subscribe - When people think about starting a software company, they usually imagine raising venture capital, hiring engineers, and growing a team as quickly as possible.
Michael Drogalis chose a different path.
After helping build technology in the Kafka ecosystem, founding a startup that was ultimately acquired by Confluent, and leading product for stream processing, he walked away from big tech to see if one person could build a serious B2B software company.
* The result became ShadowTraffic, a product that helps engineering teams generate realistic production traffic for testing, demos, and development.
Data Engineering Central is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
In this conversation, we talk about much more than streaming systems. We discuss why most engineers underestimate the importance of understanding customers, how AI is changing software development without replacing experienced engineers, what it takes to market technical products, and why writing publicly can become one of the biggest accelerators of your career.
If you’ve ever considered building your own product, becoming a solopreneur, or simply becoming a better engineer, this conversation is packed with practical advice from someone who’s actually done it.
I think this episode has broad appeal beyond data engineering. It’s really about engineering careers, entrepreneurship, and building products that solve real problems, which should make it one of your more accessible interviews.
* Building and selling a Kafka startup
* Life inside Confluent during its rapid growth
* Why Michael left big tech to become a solopreneur
* Building ShadowTraffic from scratch
* Finding customers before writing code
* Why marketing matters more than most engineers think
* Using AI without becoming dependent on it
* The future of software engineering
* Writing online and building an audience
* Advice for engineers who want to start their own business
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit dataengineeringcentral.substack.com/subscribe - The creator of Pandas and co-creator of Apache Arrow, Wes McKinney, joins the Data Engineering Central Podcast for an in-depth conversation about how modern data engineering came to exist, where AI is taking software development, and why good engineering still matters more than ever.
We start with Wes’ journey from building GoldenEye fan websites as a teenager to creating Pandas while working at a quantitative hedge fund, and eventually launching Apache Arrow, one of the foundational technologies behind today’s modern data ecosystem. Along the way, we discuss Cloudera, Parquet, DuckDB, DataFusion, Spark, and how the industry evolved from Hadoop to today’s lakehouse architectures.
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
The second half of the conversation dives deep into AI. Wes explains why large language models make experienced engineers more productive but won’t magically replace software engineering, why architecture and good taste are becoming more valuable than writing individual lines of code, and why projects like DuckDB and
* Apache Arrow remains incredibly difficult to recreate with AI alone. We also discuss open-source, local AI models, token costs, multimodal data platforms, and what new engineers should focus on to build long-term careers in software and data.
If you’re a data engineer, software engineer, architect, engineering leader, or simply interested in where AI is taking our industry, this is a conversation you won’t want to miss.
Topics We Cover
* How Pandas was created
* The story behind Apache Arrow
* Why Arrow became the standard for modern data systems
* DuckDB, DataFusion, and the next generation of data tools
* The evolution from Hadoop to lakehouses
* Why AI won’t replace great software engineers
* Architecture vs. coding in the AI era
* Building trustworthy open source software
* The future of data engineering
* Advice for new engineers entering the industry
If you enjoy conversations with the people building the future of data engineering, subscribe for more interviews with the creators of the tools we use every day.
Data Engineering Central is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit dataengineeringcentral.substack.com/subscribe - What happens when someone who started programming on a Commodore 64 watches AI reshape the entire data industry?
In this episode of the Data Engineering Central Podcast, I sit down with Dave Langer to explore nearly three decades of experience across software engineering, business intelligence, analytics, data science, and AI.
Dave’s career spans COBOL programming on IBM mainframes, enterprise architecture, Microsoft’s Xbox division, machine learning, startup leadership, authorship, and building one of the largest personal brands in the data space.
We discuss why many of the biggest problems in data haven’t changed, even as the tools continue to evolve. We dive into the reality behind self-service analytics, the importance of dimensional modeling, what organizations are getting wrong about AI adoption, and why developing strong analytical skills matters more than ever.
* Dave also shares practical advice for data professionals navigating the AI era, explaining why tools like Copilot should be viewed as partners rather than replacements.
If you’re a data analyst, BI developer, data scientist, or data engineer wondering what the future holds, this conversation offers both perspective and optimism.
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
What We Cover
* How Dave got started programming on a Commodore 64
* The transition from COBOL to modern analytics
* Why the core problems in data haven’t changed
* The evolution of business intelligence and dashboards
* How Dave discovered machine learning
* Why data science needs more than just Jupyter notebooks
* The limitations of self-service analytics
* Why semantic layers and governance matter for AI
* Advice for staying relevant as AI reshapes the industry
* Building a personal brand in data
* Writing a technical book and becoming an independent creator
Connect with Dave:
* LinkedIn: https://www.linkedin.com/in/davelanger/
* Substack: The DIY Data Scientist
* Book: Python and Excel Step-by-Step
Subscribe for more conversations on data engineering, analytics, AI, and building a career in modern data.
Data Engineering Central is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit dataengineeringcentral.substack.com/subscribe - In this episode of the Data Engineering Central Podcast, I sit down with David Jaitillake to explore the future of data engineering, analytics, and AI. David has spent nearly two decades working across data teams, from analyst roles in the early SQL Server days to leading teams, founding startups, serving as VP of AI at Cube, and now co-founding Quarry.
We discuss why semantic layers have suddenly become one of the most important concepts in modern data platforms, how tools like Claude Code are transforming engineering workflows, and why the core problems in data haven’t really changed despite massive advances in technology.
David shares his perspective on where agentic workflows are headed, what AI means for junior engineers entering the field, and why experienced practitioners may be more valuable than ever before. We also dive into the evolution of data platforms, lessons learned from startups, the promise of tools like DuckDB and MotherDuck, and how organizations should think about adopting AI responsibly.
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
Whether you’re a data engineer, analytics engineer, engineering leader, or someone trying to understand where the industry is headed, this conversation offers a practical and honest look at what’s coming next.
What We Cover
* David’s journey from analyst to startup founder
* The rise of semantic layers and why they matter
* Why data modeling is still critical in the AI era
* How AI coding agents are changing engineering work
* What Claude Code is enabling today
* The future of agentic data pipelines
* Why DuckDB and MotherDuck are gaining traction
* The challenges facing junior engineers
* Career advice for data professionals at every stage
* Whether David is optimistic about the future of AI and data
Connect with David:
* LinkedIn: https://www.linkedin.com/in/david-jayatillake/
* Substack:
Thanks for reading Data Engineering Central! This post is public so feel free to share it.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit dataengineeringcentral.substack.com/subscribe
More News podcasts
Trending News podcasts
About Data Engineering Central Podcast
Long Live the Data Engineer. No holds barred. Talking about Data Engineering news, topics, and general mayhem. dataengineeringcentral.substack.com
Podcast websiteListen to Data Engineering Central Podcast, The MeidasTouch Podcast and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


Data Engineering Central Podcast
Scan code,
download the app,
start listening.
download the app,
start listening.
























