<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>apache-spark</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/3343/feed"/>
    <updated>2026-08-26T20:39:11+00:00</updated>
    <id>https://links.biapy.com/guest/tags/3343/feed</id>
            <entry>
            <id>https://links.biapy.com/links/11443</id>
            <title type="text"><![CDATA[DLT-META]]></title>
            <link rel="alternate" href="https://databrickslabs.github.io/dlt-meta/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11443"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Metadata driven Spark Declarative Pipelines framework for bronze/silver pipelines.

DLT-META is a metadata-driven framework designed to work with Lakeflow Declarative Pipelines. This framework enables the automation of bronze and silver data pipelines by leveraging metadata recorded in an onboarding JSON file. This file, known as the Dataflowspec, serves as the data flow specification, detailing the source and target metadata required for the pipelines.

- [DLT-META @ GitHub](https://github.com/databrickslabs/dlt-meta).

Related contents:

- [From Chaos to Scale: Templatizing Spark Declarative Pipelines with DLT-META @ databricks](https://www.databricks.com/blog/chaos-scale-templatizing-spark-declarative-pipelines-dlt-meta).]]>
            </summary>
            <updated>2026-01-12T13:24:35+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11421</id>
            <title type="text"><![CDATA[Apache Spark]]></title>
            <link rel="alternate" href="https://spark.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11421"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Unified Engine for large-scale data analytics.

Apache Spark™ is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters. 

- [Apache Spark @ GitHub](https://github.com/apache/spark).

Related contents:

- [Introducing Apache Spark® 4.1 @ databricks](https://www.databricks.com/blog/introducing-apache-sparkr-41).
- [From Chaos to Scale: Templatizing Spark Declarative Pipelines with DLT-META @ databricks](https://www.databricks.com/blog/chaos-scale-templatizing-spark-declarative-pipelines-dlt-meta).
- [Breaking the Microbatch Barrier: The Architecture of Apache Spark Real-Time Mode @ databricks](https://www.databricks.com/blog/breaking-microbatch-barrier-architecture-apache-spark-real-time-mode).]]>
            </summary>
            <updated>2026-03-17T12:31:57+00:00</updated>
        </entry>
    </feed>
