Back to search

Article

Benchmarking Large Language Models for Data Pipeline Code Generation and Execution

2025-07-02

Abstract excerpt

<title>Abstract</title> <p>In today’s data-driven landscape, organizations face mounting pressure to accelerate data processing and analysis while minimizing manual engineering efforts. This paper investigates the application of Large Language Models (LLMs) to automate the creation of data pipelines, evaluating their efficacy across code-based (Apache Airflow), low-code (Azure Data Factory), and hybrid (Databrick...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
bfd9caab-abeb-5faf-b0d2-fef9f06ec8b0
DOI
10.21203/rs.3.rs-6786102/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Benchmarking Large Language Models for Data Pipeline Code Generation and ExecutionDOI 10.21203/rs.3.rs-6786102/v1
Select a neighboring publication to make it the new centre.