<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>data-lake</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/743/feed"/>
    <updated>2026-09-21T06:11:04+00:00</updated>
    <id>https://links.biapy.com/guest/tags/743/feed</id>
            <entry>
            <id>https://links.biapy.com/links/13450</id>
            <title type="text"><![CDATA[StarRocks]]></title>
            <link rel="alternate" href="https://www.starrocks.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13450"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A High-Performance Analytical Database.

The world&amp;#039;s fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.

- [StarRocks @ GitHub](https://github.com/StarRocks/StarRocks).]]>
            </summary>
            <updated>2026-07-29T13:40:47+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12586</id>
            <title type="text"><![CDATA[Project Nessie]]></title>
            <link rel="alternate" href="https://projectnessie.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12586"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Transactional Catalog for Data Lakes with Git-like semantics.

Nessie supports Iceberg Tables/Views. Additionally, Nessie is focused on working with the widest range of tools possible, which can be seen in the feature matrix.

- [Project Nessie @ GitHub](https://github.com/projectnessie/nessie/).

Related contents:

- [How to Build an Open Source Data Lake for Batch Ingestion @ freeCodeCamp](https://www.freecodecamp.org/news/how-to-build-an-open-source-data-lake-for-batch-ingestion/#heading-nessie).]]>
            </summary>
            <updated>2026-04-21T05:47:33+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11810</id>
            <title type="text"><![CDATA[SeaweedFS Enterprise]]></title>
            <link rel="alternate" href="https://seaweedfs.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11810"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Scalable Distributed Storage.

 SeaweedFS is a fast distributed storage system for blobs, objects, files, and data lake, for billions of files! Blob store has O(1) disk seek, cloud tiering. Filer supports Cloud Drive, xDC replication, Kubernetes, POSIX FUSE mount, S3 API, S3 Gateway, Hadoop, WebDAV, encryption, Erasure Coding. Enterprise version is at seaweedfs.com. 

- [SeaweedFS @ GitHub](https://github.com/seaweedfs/seaweedfs).]]>
            </summary>
            <updated>2026-02-13T13:53:07+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11631</id>
            <title type="text"><![CDATA[Microsoft Fabric]]></title>
            <link rel="alternate" href="https://www.microsoft.com/en-us/microsoft-fabric" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11631"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Data Analytics Platform.

Related conttens:

- [Qu&amp;#039;est-ce que Microsoft Fabric ? @ datacamp :fr:](https://www.datacamp.com/fr/blog/what-is-microsoft-fabric).]]>
            </summary>
            <updated>2026-01-27T11:05:00+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11443</id>
            <title type="text"><![CDATA[DLT-META]]></title>
            <link rel="alternate" href="https://databrickslabs.github.io/dlt-meta/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11443"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Metadata driven Spark Declarative Pipelines framework for bronze/silver pipelines.

DLT-META is a metadata-driven framework designed to work with Lakeflow Declarative Pipelines. This framework enables the automation of bronze and silver data pipelines by leveraging metadata recorded in an onboarding JSON file. This file, known as the Dataflowspec, serves as the data flow specification, detailing the source and target metadata required for the pipelines.

- [DLT-META @ GitHub](https://github.com/databrickslabs/dlt-meta).

Related contents:

- [From Chaos to Scale: Templatizing Spark Declarative Pipelines with DLT-META @ databricks](https://www.databricks.com/blog/chaos-scale-templatizing-spark-declarative-pipelines-dlt-meta).]]>
            </summary>
            <updated>2026-01-12T13:24:35+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11238</id>
            <title type="text"><![CDATA[mooncake]]></title>
            <link rel="alternate" href="https://www.mooncake.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11238"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[a data lakehouse for you and me. managed + real-time Iceberg.

🥮 is real-time + managed Apache Iceberg.
bringing open analytical tables on object store to every team.

pg_mooncake is a ClickHouse alternative for real-time analytics built on Postgres. It turns Postgres into a real-time analytics database by adding:

    Columnar storage (Apache Iceberg, via Moonlink)
    Vectorized execution with DuckDB (via pg_duckdb).

Fast analytics queries require both columnar storage &amp;amp; vectorized execution, and previous Postgres analytics solutions only solved half the problem.

- [mooncake @ GitHub](https://github.com/Mooncake-Labs/).]]>
            </summary>
            <updated>2025-12-15T10:10:52+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10890</id>
            <title type="text"><![CDATA[pg_lake]]></title>
            <link rel="alternate" href="https://github.com/Snowflake-Labs/pg_lake" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10890"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Postgres with Iceberg and data lake access.

pg_lake integrates Iceberg and data lake files into Postgres. With the pg_lake extensions, you can use Postgres as a stand-alone lakehouse system that supports transactions and fast queries on Iceberg tables, and can directly work with raw data files in object stores like S3.

Related contents:

- [Postgres Is the Gateway Drug @ Vignesh Ravichandran](https://viggy28.dev/article/postgres-gateway-drug/).]]>
            </summary>
            <updated>2026-03-23T16:36:25+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10258</id>
            <title type="text"><![CDATA[Bauplan]]></title>
            <link rel="alternate" href="https://www.bauplanlabs.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10258"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Your data lakehouse, built like software.

Bauplan is a cloud-native lakehouse platform for engineering teams who treat data like software.
Ship pipelines without managing infrastructure, using a specialized Python runtime, Git-for-Data built on Apache Iceberg, and just a few simple APIs.

Related contents:

- [Bauplan: Operate your lakehouse with zero infrastructure @ Data Engineer Things](https://blog.dataengineerthings.org/bauplan-operate-your-lakehouse-with-zero-infrastructure-f15a24ca33a9).]]>
            </summary>
            <updated>2025-09-18T05:55:14+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/661</id>
            <title type="text"><![CDATA[DuckLake]]></title>
            <link rel="alternate" href="https://ducklake.select/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/661"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[DuckLake is an integrated data lake and catalog format

DuckLake delivers advanced data lake features without traditional lakehouse complexity by using Parquet files and your SQL database. It&amp;#039;s an open, standalone format from the DuckDB team.

DuckLake is an open Lakehouse format that is built on SQL and Parquet. DuckLake stores metadata in a catalog database, and stores data in Parquet files. The DuckLake extension allows DuckDB to directly read and write data from DuckLake.

- [Ducklake @ GitHub](https://github.com/duckdb/ducklake).]]>
            </summary>
            <updated>2025-08-28T17:48:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/684</id>
            <title type="text"><![CDATA[BigQuery]]></title>
            <link rel="alternate" href="http://BigQuery" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/684"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AI data platform.

From data warehouse to autonomous data and AI platform

BigQuery is the autonomous data to AI platform, automating the entire data life cycle, from ingestion to AI-driven insights, so you can go from data to AI to action faster.

Gemini in BigQuery features are now included in BigQuery pricing models.

Related contents:

- [BigQuery’s Ridiculous Pricing Model Cost Us $10,000 in Just 22 Seconds!!! @ Data Engineer Things](https://blog.det.life/bigquerys-ridiculous-pricing-model-cost-us-10-000-in-just-22-seconds-7d52e3e4ae60).]]>
            </summary>
            <updated>2025-08-28T17:52:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1110</id>
            <title type="text"><![CDATA[OLake]]></title>
            <link rel="alternate" href="https://olake.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1110"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Fastest way to Replicate your Database data in Data Lake.
OLake makes data replication faster by parallelizing full loads, leveraging change streams for real-time sync, and pulling data in a database-native format for efficient ingestion.

Fastest open-source tool for replicating Databases to Apache Iceberg or Data Lakehouse. ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supporting Postgres, MongoDB and MySQL 

- [OLake @ GitHub](https://github.com/datazip-inc/olake).

Related contents:

- [Change Data Capture Tools @ Dev Genius&amp;#039; Medium](https://blog.devgenius.io/change-data-capture-tools-c0e4ee4434ac).]]>
            </summary>
            <updated>2025-08-28T19:02:48+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1357</id>
            <title type="text"><![CDATA[Apache Gravitino]]></title>
            <link rel="alternate" href="https://gravitino.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1357"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A unified metadata lake across all your sources, formats, cloud providers, and regions in a federated architecture.
 World&amp;#039;s most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.

Apache Gravitino is a high-performance, geo-distributed, and federated metadata lake. It manages metadata directly in different sources, types, and regions, providing users with unified metadata access for data and AI assets. 

- [Apache Gravitino @ GitHub](https://github.com/apache/gravitino).]]>
            </summary>
            <updated>2025-08-28T19:43:10+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3613</id>
            <title type="text"><![CDATA[Apache Hudi]]></title>
            <link rel="alternate" href="https://hudi.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3613"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[An Open Source Data Lake Platform.

Apache Hudi is a transactional data lake platform that brings database and data warehouse capabilities to the data lake. Hudi reimagines slow old-school batch data processing with a powerful new incremental processing framework for low latency minute-level analytics.

- [Apache Hudi @ GitHub](https://github.com/apache/hudi).
- [Building and scaling Notion’s data lake @ Notion blog](https://www.notion.so/blog/building-and-scaling-notions-data-lake).]]>
            </summary>
            <updated>2025-08-29T01:59:14+00:00</updated>
        </entry>
    </feed>
