Integrating a Data Warehouse and a Data Lake

Integrating a Data Warehouse and a Data Lake

TNS host Alex Williams is joined by Florian Valeye, a data engineer at Back Market, to shed light on the evolving landscape of data engineering, particularly focusing on Delta Lake and his contributions to open source communities. As a member of the Delta Lake community, Valeye discusses the intersection of data warehouses and data lakes, emphasizing the need for a unified platform that breaks down traditional barriers.

Delta Lake, initially created by Databricks and now under the Linux Foundation, aims to enhance reliability, performance, and quality in data lakes. Valeye explains how Delta Lake addresses the challenges posed by the separation of data warehouses and data lakes, emphasizing the importance of providing asset transactions, real-time processing, and scalable metadata.

Valeye's involvement in Delta Lake began as a response to the challenges faced at Back Market, a global marketplace for refurbished devices. The platform manages large datasets, and Delta Lake proved to be a pivotal solution in optimizing ETL processes and facilitating communication between data scientists and data engineers.

The conversation delves into Valeye's journey with Delta Lake, his introduction to Rust programming language, and his role as a maintainer in the Rust-based library for Delta Lake. Valeye emphasizes Rust's importance in providing a high-level API with reliability and efficiency, offering a balanced approach for developers.

Looking ahead, Valeye envisions Delta Lake evolving beyond traditional data engineering, becoming a platform that seamlessly connects data scientists and engineers. He anticipates improvements in data storage optimization and envisions Delta Lake serving as a standard format for machine learning and AI applications.

The conversation concludes with Valeye reflecting on his future contributions, expressing a passion for Rust programming and an eagerness to explore evolving projects in the open-source community.

Learn more from The New Stack about Delta Lake and The Linux Foundation:

Delta Lake: A Layer to Ensure Data Quality

Data in 2023: Revenge of the SQL Nerds

What Do You Know about Your Linux System?

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(300)

Dynatrace's $915M Arize deal bets AI agents are just another app to monitor

Dynatrace's $915M Arize deal bets AI agents are just another app to monitor

Dynatrace completed its $915 million acquisition of Arize on Oct. 1, combining Dynatrace’s application and infrastructure monitoring with Arize’s AI agent tracing and evaluation capabilities. The deal...

5 Okt 20min

Bit Cloud’s next chapter starts after the AI builds your app

Bit Cloud’s next chapter starts after the AI builds your app

For many developers, turning an AI-generated prototype into maintainable software requires more than generating code—it requires infrastructure, collaboration, testing and review. In this episode of T...

1 Okt 31min

CloudBees just committed to an AI-first pivot. Here's why it matters for enterprise DevOps teams

CloudBees just committed to an AI-first pivot. Here's why it matters for enterprise DevOps teams

CloudBees CEO Mo Plassnig is leading the CI/CD company through a major transformation as generative AI reshapes software development. Returning to CloudBees eight years after joining through its acqui...

30 Sep 23min

A third option is emerging in the fight over AI and your data

A third option is emerging in the fight over AI and your data

The AI industry has faced a growing enterprise dilemma: companies want access to powerful proprietary AI models without risking sensitive data or intellectual property, while AI labs want to protect t...

23 Sep 30min

Drowning in AI pull requests: Harness's field CTO on code review and a Git repo built for agents

Drowning in AI pull requests: Harness's field CTO on code review and a Git repo built for agents

Harness Field CTO Martin Reynolds joins The New Stack to talk about what happens after coding agents start opening pull requests faster than anyone can review them. He explains how he first saw the bo...

7 Sep 26min

How to find failures without drowning in tracing data

How to find failures without drowning in tracing data

Traces provide a detailed view of a request’s journey through data, microservices and applications, helping SREs pinpoint where failures occur and resolve issues faster. But while tracing can reduce d...

3 Sep 31min

Why CPUs still matter in the age of AI agents

Why CPUs still matter in the age of AI agents

As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Pate...

11 Aug 26min

Why Doist Says Less AI Can Deliver More

Why Doist Says Less AI Can Deliver More

Doist CTO Gonçalo Silva says AI is reshaping software development, but success depends on restraint rather than rapid feature expansion. Instead of chasing every AI capability, Doist prioritizes “subt...

31 Jul 35min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden
aftenpodden-usa
forklart
popradet
fotballpodden-2
stopp-verden
det-store-bildet
dine-penger-pengeradet
rss-espen-lee-usensurert
nokon-ma-ga
rss-gukild-johaug
hanna-de-heldige
aftenbla-bla
rss-ness
frokostshowet-pa-p5
bt-dokumentar-2
e24-podden
rss-penger-polser-og-politikk
ta-dokumentar