<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>big-data</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/684/feed"/>
    <updated>2026-08-02T22:09:10+00:00</updated>
    <id>https://links.biapy.com/guest/tags/684/feed</id>
            <entry>
            <id>https://links.biapy.com/links/13455</id>
            <title type="text"><![CDATA[Semarchy]]></title>
            <link rel="alternate" href="https://semarchy.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13455"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[native snowflake MDM.]]>
            </summary>
            <updated>2026-07-29T15:32:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13452</id>
            <title type="text"><![CDATA[Snowflake]]></title>
            <link rel="alternate" href="https://www.snowflake.com/en/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13452"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Snowflake is a cloud-based data warehousing platform designed to provide a flexible, scalable, and high-performance solution for storing, processing, and analyzing large volumes of data. It leverages a unique architecture that separates storage and compute resources, allowing users to scale each independently.]]>
            </summary>
            <updated>2026-07-29T14:22:23+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11421</id>
            <title type="text"><![CDATA[Apache Spark]]></title>
            <link rel="alternate" href="https://spark.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11421"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Unified Engine for large-scale data analytics.

Apache Spark™ is a multi-language engine for executing data engineering, data science, and machine learning on single-node machines or clusters. 

- [Apache Spark @ GitHub](https://github.com/apache/spark).

Related contents:

- [Introducing Apache Spark® 4.1 @ databricks](https://www.databricks.com/blog/introducing-apache-sparkr-41).
- [From Chaos to Scale: Templatizing Spark Declarative Pipelines with DLT-META @ databricks](https://www.databricks.com/blog/chaos-scale-templatizing-spark-declarative-pipelines-dlt-meta).
- [Breaking the Microbatch Barrier: The Architecture of Apache Spark Real-Time Mode @ databricks](https://www.databricks.com/blog/breaking-microbatch-barrier-architecture-apache-spark-real-time-mode).]]>
            </summary>
            <updated>2026-03-17T12:31:57+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10808</id>
            <title type="text"><![CDATA[Apache Calcite]]></title>
            <link rel="alternate" href="https://calcite.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10808"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Dynamic data management framework.
The foundation for your next high-performance database.

It contains many of the pieces that comprise a typical database management system but omits the storage primitives. It provides an industry standard SQL parser and validator, a customisable optimizer with pluggable rules and cost functions, logical and physical algebraic operators, various transformation algorithms from SQL to algebra (and the opposite), and many adapters for executing SQL queries over Cassandra, Druid, Elasticsearch, MongoDB, Kafka, and others, with minimal configuration.

- [Apache Calcite @ GitHub](https://github.com/apache/calcite).

Related contents:

- [Reimagining log analytics for the modern enterprise @ OpenSearch](https://opensearch.org/blog/reimagining-log-analytics-for-the-modern-enterprise/).]]>
            </summary>
            <updated>2025-10-29T12:43:17+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10721</id>
            <title type="text"><![CDATA[Code-First CDC to ClickHouse with Debezium, Redpanda, and MooseStack]]></title>
            <link rel="alternate" href="https://github.com/514-labs/debezium-cdc" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10721"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Easy-to-run demo of a CDC pipeline using Debezium (Kafka Connect), PostgreSQL, Redpanda, and ClickHouse.]]>
            </summary>
            <updated>2025-10-20T06:22:14+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/569</id>
            <title type="text"><![CDATA[Datafari Enterprise Search]]></title>
            <link rel="alternate" href="https://www.datafari.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/569"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Source, Distributed, Big Data Enterprise Search Engine.

Datafari is an open source enterprise search solution enriched with AI. It is the perfect product for anyone who needs to search and analyze its corporate data and documents, both within the content and the metadata. Plus, with its genAI modules, it allows to easily leverage mistral, openai, or local LLMs for your company data.

- [Datafari @ GitHub](https://github.com/francelabs/datafari).]]>
            </summary>
            <updated>2025-08-28T17:34:01+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/684</id>
            <title type="text"><![CDATA[BigQuery]]></title>
            <link rel="alternate" href="http://BigQuery" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/684"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AI data platform.

From data warehouse to autonomous data and AI platform

BigQuery is the autonomous data to AI platform, automating the entire data life cycle, from ingestion to AI-driven insights, so you can go from data to AI to action faster.

Gemini in BigQuery features are now included in BigQuery pricing models.

Related contents:

- [BigQuery’s Ridiculous Pricing Model Cost Us $10,000 in Just 22 Seconds!!! @ Data Engineer Things](https://blog.det.life/bigquerys-ridiculous-pricing-model-cost-us-10-000-in-just-22-seconds-7d52e3e4ae60).]]>
            </summary>
            <updated>2025-08-28T17:52:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1153</id>
            <title type="text"><![CDATA[Apache Kafka]]></title>
            <link rel="alternate" href="https://kafka.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1153"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Apache Kafka is an open-source distributed event streaming platform used by thousands of companies for high-performance data pipelines, streaming analytics, data integration, and mission-critical applications.

Related contents:

- [Why Was Apache Kafka Created? @ Big Data Stream](https://bigdata.2minutestreaming.com/p/why-was-apache-kafka-created).

- [Apache Kafka @ GitHub](https://github.com/apache/kafka).

Related contents:

- [The New Look and Feel of Apache Kafka 4.0 @ The New Stack](https://thenewstack.io/the-new-look-and-feel-of-apache-kafka-4-0/).
- [Kafka: The End of the Beginning @ Materialized View](https://materializedview.io/p/kafka-end-of-beginning).
- [Optimizing Kafka Tracing with OpenTelemetry: Boost Visibility &amp;amp; Performance @ New Relic](https://newrelic.com/blog/how-to-relic/optimizing-kafka-tracing-with-opentelemetry-boost-visibility-performance).
- [Introducing Apache Kafka® 4.1.0: What’s New and How to Upgrade @ Confluent](https://www.confluent.io/blog/introducing-apache-kafka-4-1/).
- [Testing Kafka-based Asynchronous Workflows Using OpenTelemetry @ Signadot](https://www.signadot.com/blog/testing-kafka-based-asynchronous-workflows-using-opentelemetry).
- [Kafka is fast -- I&amp;#039;ll use Postgres @ TopicPartition](https://topicpartition.io/blog/postgres-pubsub-queue-benchmarks).
- [The KFC Architecture Blueprint: Kafka, Flink, and ClickHouse @ Big Data Boutique](https://bigdataboutique.com/blog/kfc-architecture-blueprint-kafka-flink-and-clickhouse).
- [Episode #147 - RabbitMQ, Kafka et les messages brokers @ Code Garage :fr:](https://code-garage.com/podcast/classic/episode-147).]]>
            </summary>
            <updated>2026-03-19T07:15:38+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1357</id>
            <title type="text"><![CDATA[Apache Gravitino]]></title>
            <link rel="alternate" href="https://gravitino.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1357"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A unified metadata lake across all your sources, formats, cloud providers, and regions in a federated architecture.
 World&amp;#039;s most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.

Apache Gravitino is a high-performance, geo-distributed, and federated metadata lake. It manages metadata directly in different sources, types, and regions, providing users with unified metadata access for data and AI assets. 

- [Apache Gravitino @ GitHub](https://github.com/apache/gravitino).]]>
            </summary>
            <updated>2025-08-28T19:43:10+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1949</id>
            <title type="text"><![CDATA[Apache Pinot™]]></title>
            <link rel="alternate" href="https://pinot.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1949"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Insights, Unlocked in Real Time.

Apache Pinot™: The real-time analytics open source platform for lightning-fast insights, effortless scaling, and cost-effective data-driven decisions.

- [Apache Pinot @ GitHub](https://github.com/apache/pinot).

Related contents:

- [Serving Millions of Apache Pinot™ Queries with Neutrino @ Uber Blog](https://www.uber.com/en-FR/blog/serving-millions-of-apache-pinot-queries-with-neutrino/?uclick_id=ed80e0fe-d305-48c1-b7e9-ed149ec25b99).]]>
            </summary>
            <updated>2025-08-28T21:21:06+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2593</id>
            <title type="text"><![CDATA[Apache Iceberg™]]></title>
            <link rel="alternate" href="https://iceberg.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2593"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The open table format for analytic datasets.

Iceberg is a high-performance format for huge analytic tables. Iceberg brings the reliability and simplicity of SQL tables to big data, while making it possible for engines like Spark, Trino, Flink, Presto, Hive and Impala to safely work with the same tables, at the same time.

- [Apache Iceberg @ GitHub](https://github.com/apache/iceberg).

Related contents:

- [PyIceberg: Current State and Roadmap @ Ju Data Engineering Newsletter](https://juhache.substack.com/p/pyiceberg-current-state-and-roadmap).
- [The Equality Delete Problem in Apache Iceberg @ Data Engineer Things&amp;#039;s Medium](https://blog.dataengineerthings.org/the-equality-delete-problem-in-apache-iceberg-143dd451a974).
- [How I Saved Millions by Restructuring Iceberg Metadata @ Gautham Gondi&amp;#039;s Medium](https://medium.com/@gauthamnagendra/how-i-saved-millions-by-restructuring-iceberg-metadata-c4f5c1de69c2).
- [High Throughput Ingestion with Iceberg @ Adobe Tech Blog&amp;#039;s Medium](https://medium.com/adobetech/high-throughput-ingestion-with-iceberg-ccf7877a413f).
- [Scaling Iceberg Writes with Confidence: A Conflict-Free Distributed Architecture for Fast, Concurrent, Consistent Append-Only Writes @ e6data](https://www.e6data.com/blog/iceberg-distributed-architecture-fast-concurrent-append-writes).
- [Postgres Is the Gateway Drug @ Vignesh Ravichandran](https://viggy28.dev/article/postgres-gateway-drug/).]]>
            </summary>
            <updated>2026-03-23T16:38:12+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2868</id>
            <title type="text"><![CDATA[Bufstream]]></title>
            <link rel="alternate" href="https://buf.build/product/bufstream" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2868"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The best way of working with Protocol Buffers.  Elastic, self-hosted Kafka with Advanced Semantic Intelligence
Guarantee streaming data quality and slash cloud costs 10x with Bufstream, a drop-in replacement for Apache Kafka®.

Bufstream is a Kafka-compatible streaming system which stores records directly in an object storage service like S3.

- [Buf @ GitHub](https://github.com/bufbuild/buf).
- [Bufstream 0.1.0 @ Jebsen](https://jepsen.io/analyses/bufstream-0.1.0).]]>
            </summary>
            <updated>2025-08-28T23:54:30+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3013</id>
            <title type="text"><![CDATA[Apache CouchDB]]></title>
            <link rel="alternate" href="https://couchdb.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3013"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Seamless multi-master sync, that scales from Big Data to Mobile, with an Intuitive HTTP/JSON API and designed for Reliability.

CouchDB is a database that completely embraces the web. Store your data with JSON documents. Access your documents with your web browser, via HTTP. Query, combine, and transform your documents with JavaScript. CouchDB works well with modern web and mobile apps. You can distribute your data, efficiently using CouchDB’s incremental replication. CouchDB supports master-master setups with automatic conflict detection.

Related contents:

- [How to Sync Anything @ Neighbourhoodie Software](https://neighbourhood.ie/blog/2025/04/06/how-to-sync-anything).
- [Offline-First with CouchDB and PouchDB in 2025 @ Neighbourhoodie Software](https://neighbourhood.ie/blog/2025/03/26/offline-first-with-couchdb-and-pouchdb-in-2025).]]>
            </summary>
            <updated>2025-08-29T00:18:17+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3062</id>
            <title type="text"><![CDATA[Trench]]></title>
            <link rel="alternate" href="https://www.trench.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3062"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open source analytics infrastructure. Fast and scalable. No bloat. GDPR compliant.

A single production-ready Docker image built on ClickHouse, Kafka, and Node.js for tracking events, users, page views, and interactions. 

- [Trench @ GitHub](https://github.com/FrigadeHQ/trench).]]>
            </summary>
            <updated>2025-08-29T00:28:23+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3303</id>
            <title type="text"><![CDATA[Apache Hadoop]]></title>
            <link rel="alternate" href="https://hadoop.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3303"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The Apache® Hadoop® project develops open-source software for reliable, scalable, distributed computing.

The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.

- [Apache Hadoop @ GitHub](https://github.com/apache/hadoop).

Related contents:

- [Why NoSQL Deployments Are Failing at Scale @ The New Stack](https://thenewstack.io/why-nosql-deployments-are-failing-at-scale/).
- [From SSH to REST: A Security-Driven Modernization of Slack’s EMR Data Pipelines @ slack engineering](https://slack.engineering/from-ssh-to-rest-a-security-driven-modernization-of-slacks-emr-data-pipelines/).]]>
            </summary>
            <updated>2026-05-15T13:44:51+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3487</id>
            <title type="text"><![CDATA[AutoMQ]]></title>
            <link rel="alternate" href="https://www.automq.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3487"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Source Available Reinvented Kafka. 10x Cost Efficiency.

AutoMQ is a cloud-first alternative to Kafka by decoupling durability to S3 and EBS. 10x cost-effective. Autoscale in seconds. Single-digit ms latency. 

 AutoMQ is a stateless Kafka on S3. 10x Cost-Effective. No Cross-AZ Traffic Cost. Autoscale in seconds. Single-digit ms latency. Multi-AZ Availability. 

- [AutoMQ @ GitHub](https://github.com/AutoMQ/automq).

Related contents:

- [What If We Could Rebuild Kafka From Scratch? @ Gunnar Morling](https://www.morling.dev/blog/what-if-we-could-rebuild-kafka-from-scratch/).]]>
            </summary>
            <updated>2025-08-29T01:37:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3613</id>
            <title type="text"><![CDATA[Apache Hudi]]></title>
            <link rel="alternate" href="https://hudi.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3613"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[An Open Source Data Lake Platform.

Apache Hudi is a transactional data lake platform that brings database and data warehouse capabilities to the data lake. Hudi reimagines slow old-school batch data processing with a powerful new incremental processing framework for low latency minute-level analytics.

- [Apache Hudi @ GitHub](https://github.com/apache/hudi).
- [Building and scaling Notion’s data lake @ Notion blog](https://www.notion.so/blog/building-and-scaling-notions-data-lake).]]>
            </summary>
            <updated>2025-08-29T01:59:14+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3739</id>
            <title type="text"><![CDATA[Data For Good]]></title>
            <link rel="alternate" href="https://dataforgood.fr/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3739"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Les technologies numériques sont incroyablement puissantes et redéfinissent le fonctionnement de notre société. Pour les acteurs qui œuvrent pour l&amp;#039;intérêt général, la technologie peut parfois être un levier démutiplicateur d&amp;#039;impacts positifs, cependant et malheureusement ces acteurs n&amp;#039;ont souvent pas les ressources technologiques ou humaines pour accélérer leur action citoyenne. Data for Good existe pour rétablir l&amp;#039;équilibre.

- [286 - Data &amp;amp; Dev - Christophe Blefari @ &amp;lt;ifttd&amp;gt; :fr:](https://www.ifttd.io/episodes/data-dev).
- [289 - Data 4 Good - Ronan Sy @ &amp;lt;ifttd&amp;gt; :fr:](https://www.ifttd.io/episodes/data-4-good).]]>
            </summary>
            <updated>2025-08-29T02:19:28+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3741</id>
            <title type="text"><![CDATA[Amazon Athena]]></title>
            <link rel="alternate" href="https://aws.amazon.com/athena/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3741"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Interactive SQL. Analyze petabyte-scale data where it lives with ease and flexibility.

Amazon Athena is a serverless, interactive analytics service built on open-source frameworks, supporting open-table and file formats. Athena provides a simplified, flexible way to analyze petabytes of data where it lives. Analyze data or build applications from an Amazon Simple Storage Service (S3) data lake and 30 data sources, including on-premises data sources or other cloud systems using SQL or Python. Athena is built on open-source Trino and Presto engines and Apache Spark frameworks, with no provisioning or configuration effort required.

- [286 - Data &amp;amp; Dev - Christophe Blefari @ &amp;lt;ifttd&amp;gt; :fr:](https://www.ifttd.io/episodes/data-dev).]]>
            </summary>
            <updated>2025-08-29T02:21:28+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4018</id>
            <title type="text"><![CDATA[Apache DataFusion]]></title>
            <link rel="alternate" href="https://datafusion.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4018"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[DataFusion is a very fast, extensible query engine for building high-quality data-centric systems in Rust, using the Apache Arrow in-memory format.

DataFusion is great for building projects such as domain specific query engines, new database platforms and data pipelines, query languages and more. It lets you start quickly from a fully working engine, and then customize those features specific to your use.

- [DataFusion @ GitHub](https://github.com/apache/datafusion).]]>
            </summary>
            <updated>2025-08-29T03:06:00+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4206</id>
            <title type="text"><![CDATA[JuiceFS]]></title>
            <link rel="alternate" href="https://juicefs.com/en/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4206"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Source Distributed POSIX File System for Cloud. JuiceFS is a distributed POSIX file system built on top of Redis and S3. 

JuiceFS is a high-performance POSIX file system released under Apache License 2.0, particularly designed for the cloud-native environment. The data, stored via JuiceFS, will be persisted in Object Storage (e.g. Amazon S3), and the corresponding metadata can be persisted in various compatible database engines such as Redis, MySQL, and TiKV based on the scenarios and requirements.

With JuiceFS, massive cloud storage can be directly connected to big data, machine learning, artificial intelligence, and various application platforms in production environments. Without modifying code, the massive cloud storage can be used as efficiently as local storage.

- [JuiceFS @ GitHub](https://github.com/juicedata/juicefs).]]>
            </summary>
            <updated>2026-05-15T15:23:29+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4744</id>
            <title type="text"><![CDATA[Trunk Data Platform (TDP)]]></title>
            <link rel="alternate" href="https://www.trunkdataplatform.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4744"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[open source big data platform.

Trunk Data Platform is an Open Source, free, Hadoop distribution.

- [Trunk Data Platform (TDP) @ GitHub](https://github.com/TOSIT-IO/TDP).
- [TDP Installation Guide @ Alliage](https://www.alliage.io/en/academy/getting-started/install).
- [Installation Guide to TDP, the 100% open source big data platform @ Adaltas](https://www.adaltas.com/en/2023/10/18/tdp-guide-getting-started/).]]>
            </summary>
            <updated>2025-08-29T05:07:04+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5040</id>
            <title type="text"><![CDATA[Apache Druid]]></title>
            <link rel="alternate" href="https://druid.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5040"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Druid is a high performance, real-time analytics database that delivers sub-second queries on streaming and batch data at scale and under load.]]>
            </summary>
            <updated>2025-08-29T05:57:18+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6185</id>
            <title type="text"><![CDATA[XetHub: fast, frictionless collaboration at scale]]></title>
            <link rel="alternate" href="https://xethub.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6185"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[XetHub brings speedy access and Git-based collaboration to large scale repositories of data, code, or any combination of files.
Our instant mount feature makes it possible to access GBs and TBs of data in seconds at the speed of localhost, while our de-duplication algorithm stores data and differences efficiently to save money and speed up development cycles.
XetHub is ideal for teams who already use Git to track their code changes, and want to leverage the power of infinite history, pull requests, and difference-based tracking for larger assets such as datasets or media files. Managing complete projects with familiar Git semantics makes change tracking and continuous integration a breeze, especially for workflows that use code to generate or augment assets.]]>
            </summary>
            <updated>2025-08-29T09:08:59+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6652</id>
            <title type="text"><![CDATA[Neo4j]]></title>
            <link rel="alternate" href="https://neo4j.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6652"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Graph Database Management System.
Neo4j Graph Data Platform. Blazing-Fast Graph, Petabyte Scale.
With proven trillion+ entity performance, developers, data scientists, and enterprises rely on Neo4j as the top choice for high-performance, scalable analytics, intelligent app development, and advanced AI/ML pipelines.

- [Neo4j @ GitHub](https://github.com/neo4j/neo4j).
 
Related contents:

- [Episode #18: NoSQL Smackdown! @ Changelog Interviews](https://changelog.com/podcast/18).]]>
            </summary>
            <updated>2025-12-09T09:25:32+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6704</id>
            <title type="text"><![CDATA[ClickHouse]]></title>
            <link rel="alternate" href="https://github.com/ClickHouse/ClickHouse" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6704"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[ClickHouse® is a free analytics DBMS for big data.
ClickHouse® is an open-source column-oriented database management system that allows generating analytical data reports in real-time.]]>
            </summary>
            <updated>2025-08-29T10:34:46+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6719</id>
            <title type="text"><![CDATA[Konbert]]></title>
            <link rel="alternate" href="https://konbert.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6719"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open big JSON, CSV Files: Online Viewer, Explorer and Converter.
View and convert big data files.
View large or small files right in your browser and export them in any format.]]>
            </summary>
            <updated>2025-08-29T10:37:51+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7098</id>
            <title type="text"><![CDATA[Planet]]></title>
            <link rel="alternate" href="https://www.planet.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7098"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Daily Earth Data to See Change and Make Better Decisions.
Planet provides daily satellite data that helps businesses, governments, researchers, and journalists understand the physical world and take action.]]>
            </summary>
            <updated>2025-08-29T11:40:20+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7102</id>
            <title type="text"><![CDATA[Climate TRACE]]></title>
            <link rel="alternate" href="https://www.climatetrace.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7102"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Climate TRACE was built to collect and share greenhouse gas emissions from anthropogenic (human) activities to facilitate climate action .]]>
            </summary>
            <updated>2025-08-29T11:42:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7104</id>
            <title type="text"><![CDATA[Robtex]]></title>
            <link rel="alternate" href="https://www.robtex.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7104"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Robtex is used for various kinds of research of IP numbers, Domain names, etc.
Robtex uses various sources to gather public information about IP numbers, domain names, host names, Autonomous systems, routes etc. It then indexes the data in a big database and provide free access to the data.
We aim to make the fastest and most comprehensive free DNS lookup tool on the Internet.
Our database now contains billions of documents of internet data collected over more than a decade.]]>
            </summary>
            <updated>2025-08-29T11:42:31+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7259</id>
            <title type="text"><![CDATA[Luna]]></title>
            <link rel="alternate" href="https://www.luna-lang.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7259"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A WYSIWYG language for data processing.]]>
            </summary>
            <updated>2025-08-29T12:07:32+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7422</id>
            <title type="text"><![CDATA[EveryPolitician]]></title>
            <link rel="alternate" href="http://everypolitician.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7422"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Political data for 233 countries.
The world’s richest open dataset on politicians]]>
            </summary>
            <updated>2025-08-29T12:34:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7602</id>
            <title type="text"><![CDATA[PNDA]]></title>
            <link rel="alternate" href="http://pndaproject.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7602"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The scalable, open source
big data analytics platform
for networks and services.]]>
            </summary>
            <updated>2025-08-29T13:05:05+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7640</id>
            <title type="text"><![CDATA[ROOT a Data analysis Framework]]></title>
            <link rel="alternate" href="https://root.cern.ch/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7640"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A modular scientific software framework. It provides all the functionalities needed to deal with big data processing, statistical analysis, visualisation and storage. It is mainly written in C++ but integrated with other languages such as Python and R.]]>
            </summary>
            <updated>2025-08-29T13:11:16+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7673</id>
            <title type="text"><![CDATA[OpenGrid]]></title>
            <link rel="alternate" href="https://github.com/Chicago/opengrid" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7673"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A user-friendly, map-based tool to combine and explore real-time or historical data.]]>
            </summary>
            <updated>2025-08-29T13:16:14+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7702</id>
            <title type="text"><![CDATA[WhereHows]]></title>
            <link rel="alternate" href="https://github.com/linkedin/WhereHows" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7702"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[WhereHows is a data discovery and lineage tool built at LinkedIn. It integrates with all the major data processing systems and collects both catalog and operational metadata from them.]]>
            </summary>
            <updated>2025-08-29T13:21:18+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7712</id>
            <title type="text"><![CDATA[GridDB]]></title>
            <link rel="alternate" href="https://github.com/griddb/griddb_nosql" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7712"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[high performance, high scalability and high reliability database for big data.
GridDB has a KVS (Key-Value Store)-type data model that is suitable for sensor data stored in a timeseries. It is a database that can be easily scaled-out according to the number of sensors.]]>
            </summary>
            <updated>2025-08-29T13:23:22+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8352</id>
            <title type="text"><![CDATA[HBase - Apache HBase&amp;amp;#153; Home]]></title>
            <link rel="alternate" href="http://hbase.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8352"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Apache HBase™ is the Hadoop database, a distributed, scalable, big data store. 
Use Apache HBase when you need random, realtime read/write access to your Big Data. This project&amp;#039;s goal is the hosting of very large tables -- billions of rows X millions of columns -- atop clusters of commodity hardware. Apache HBase is an open-source, distributed, versioned, non-relational database modeled after Google&amp;#039;s Bigtable: A Distributed Storage System for Structured Data by Chang et al. Just as Bigtable leverages the distributed data storage provided by the Google File System, Apache HBase provides Bigtable-like capabilities on top of Hadoop and HDFS.]]>
            </summary>
            <updated>2025-08-29T15:09:22+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8357</id>
            <title type="text"><![CDATA[D3.js - Data-Driven Documents]]></title>
            <link rel="alternate" href="http://d3js.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8357"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[D3.js is a JavaScript library for manipulating documents based on data. D3 helps you bring data to life using HTML, SVG and CSS. D3’s emphasis on web standards gives you the full capabilities of modern browsers without tying yourself to a proprietary framework, combining powerful visualization components and a data-driven approach to DOM manipulation.]]>
            </summary>
            <updated>2025-08-29T15:10:20+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/9000</id>
            <title type="text"><![CDATA[Waarp]]></title>
            <link rel="alternate" href="https://github.com/Waarp/Waarp-All" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/9000"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Waarp provides a secure and efficient open source MFT solution.

Waarp Platform is a set of applications and tools specialized in managing and monitoring a high number of transfers in a secure and reliable way.

It relies on its own open protocol named R66, which has been designed to optimize file transfers, ensure the integrity of the data provide ways to integrate transfers in larger business transactions.]]>
            </summary>
            <updated>2025-08-29T16:57:35+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/9148</id>
            <title type="text"><![CDATA[Apache&amp;amp;#8482; Hadoop&amp;amp;#8482;!]]></title>
            <link rel="alternate" href="http://hadoop.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/9148"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using a simple programming model. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-avaiability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-availabile service on top of a cluster of computers, each of which may be prone to failures.]]>
            </summary>
            <updated>2025-08-29T17:22:31+00:00</updated>
        </entry>
    </feed>
