<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>crawler</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/362/feed"/>
    <updated>2026-08-03T00:28:10+00:00</updated>
    <id>https://links.biapy.com/guest/tags/362/feed</id>
            <entry>
            <id>https://links.biapy.com/links/12984</id>
            <title type="text"><![CDATA[Reader]]></title>
            <link rel="alternate" href="https://reader.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12984"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AI Web Infrastructure Platform.

 Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents. 

- [Reader @ GitHub](https://github.com/vakra-dev/reader).

Related contents:

- [Top 19 des alternatives open-source aux SaaS en 2026 @ Camille Roux :fr:](https://www.camilleroux.com/alternatives-open-source-saas-2026-v2/).]]>
            </summary>
            <updated>2026-06-10T12:46:46+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12867</id>
            <title type="text"><![CDATA[Crawl4AI]]></title>
            <link rel="alternate" href="https://docs.crawl4ai.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12867"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[pen-source LLM Friendly Web Crawler &amp;amp; Scraper.

Crawl4AI turns the web into clean, LLM ready Markdown for RAG, agents, and data pipelines. Fast, controllable, battle tested by a 50k+ star community.

- [Crawl4AI @ GitHub](https://github.com/unclecode/crawl4ai).]]>
            </summary>
            <updated>2026-06-01T12:31:19+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11557</id>
            <title type="text"><![CDATA[wxpath]]></title>
            <link rel="alternate" href="https://github.com/rodricios/wxpath" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11557"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[declarative web crawling with XPath.

wxpath is a declarative web crawler where traversal is expressed directly in XPath. Instead of writing imperative crawl loops, wxpath lets you describe what to follow and what to extract in a single expression. wxpath executes that expression concurrently, breadth-first-ish, and streams results as they are discovered.]]>
            </summary>
            <updated>2026-01-21T13:06:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11052</id>
            <title type="text"><![CDATA[LibreCrawl]]></title>
            <link rel="alternate" href="https://librecrawl.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11052"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A powerful, free SEO crawler with features that surpass even paid Screaming Frog.

 Free desktop SEO crawler - open source alternative to Screaming Frog and similar tools. Crawl websites, analyze links, extract SEO data, and export results without subscription fees. Fully customizable and extensible! 

- [LibreCrawl @ GitHub](https://github.com/PhialsBasement/LibreCrawl).]]>
            </summary>
            <updated>2025-11-24T08:06:48+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10380</id>
            <title type="text"><![CDATA[SiteOne Crawler]]></title>
            <link rel="alternate" href="https://crawler.siteone.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10380"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[free website analyzer, offline exporter, sitemap generator and Swiss Army Knife, you will love.

 SiteOne Crawler is a cross-platform website crawler and analyzer for SEO, security, accessibility, and performance optimization—ideal for developers, DevOps, QA engineers, and consultants. Supports Windows, macOS, and Linux (x64 and arm64). 

- [SiteOne Crawler @ GitHub](https://github.com/janreges/siteone-crawler).]]>
            </summary>
            <updated>2025-09-24T14:14:48+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10379</id>
            <title type="text"><![CDATA[SEOnaut]]></title>
            <link rel="alternate" href="https://seonaut.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10379"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Source SEO audit tool.

SEOnaut is an SEO tool for website audits under the MIT license, giving you full transparency and control. Customize the tool to fit your unique needs or contribute to its ongoing development. Flexible, adaptable software you can trust.

- [SEOnaut @ GitHub](https://github.com/StJudeWasHere/seonaut).]]>
            </summary>
            <updated>2025-09-24T14:13:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/193</id>
            <title type="text"><![CDATA[A Vocabulary For Expressing AI Usage Preferences]]></title>
            <link rel="alternate" href="https://ietf-wg-aipref.github.io/drafts/draft-ietf-aipref-vocab.html?cf_target_id=_blank" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/193"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[This document proposes a standardized vocabulary for expressing preferences related to how digital assets are used by automated processing systems. This vocabulary allows for the creation of structured declarations about restrictions or permissions for use of digital assets by such systems.

Related contents:

- [Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives @ Cloudflare Blog](https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/).]]>
            </summary>
            <updated>2025-09-18T15:26:23+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/352</id>
            <title type="text"><![CDATA[The Web Robots Pages]]></title>
            <link rel="alternate" href="https://www.robotstxt.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/352"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Web Robots (also known as Web Wanderers, Crawlers, or Spiders), are programs that traverse the Web automatically. Search engines such as Google use them to index the web content, spammers use them to scan for email addresses, and they have many other uses.

Related contents:

- [I was wrong about robots.txt @ Evgenii Pendragon](https://evgeniipendragon.com/posts/i-was-wrong-about-robots-txt/).
- [Fix Your robots.txt or Your Site Disappears from Google @ alanwsmith.com](https://www.alanwsmith.com/en/37/wa/jz/s1/).]]>
            </summary>
            <updated>2026-01-22T12:47:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1443</id>
            <title type="text"><![CDATA[Photon]]></title>
            <link rel="alternate" href="https://github.com/s0md3v/Photon" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1443"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Incredibly fast crawler designed for OSINT.]]>
            </summary>
            <updated>2025-08-28T19:56:17+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1827</id>
            <title type="text"><![CDATA[Nepenthes]]></title>
            <link rel="alternate" href="https://zadzmo.org/code/nepenthes/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1827"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[This is a tarpit intended to catch web crawlers. Specifically, it&amp;#039;s targetting crawlers that scrape data for LLM&amp;#039;s - but really, like the plants it is named after, it&amp;#039;ll eat just about anything that finds it&amp;#039;s way inside.

It works by generating an endless sequences of pages, each of which with dozens of links, that simply go back into a the tarpit. Pages are randomly generated, but in a deterministic way, causing them to appear to be flat files that never change. Intentional delay is added to prevent crawlers from bogging down your server, in addition to wasting their time. Lastly, optional Markov-babble can be added to the pages, to give the crawlers something to scrape up and train their LLMs on, hopefully accelerating model collapse.

Related contents:

- [Nepenthes - Piégez les crawlers web malveillants @ Korben :fr:](https://korben.info/nepenthes-piege-crawlers-web-malveillants.html).
- [Open source devs are fighting AI crawlers with cleverness and vengeance @ TechCrunch](https://techcrunch.com/2025/03/27/open-source-devs-are-fighting-ai-crawlers-with-cleverness-and-vengeance/).
- [Ask HN: How to stop an AWS bot sending 2B requests/month? @ Hacker News](https://news.ycombinator.com/item?id=45613567).]]>
            </summary>
            <updated>2025-10-20T06:35:24+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1835</id>
            <title type="text"><![CDATA[Common Crawl]]></title>
            <link rel="alternate" href="https://commoncrawl.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1835"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Repository of Web Crawl Data.

Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.

Related contents:

- [S5E7 - Sommes-nous à l&amp;#039;aube d&amp;#039;un effondrement des IA ? @ Underscore_&amp;#039;s acast :fr:](https://shows.acast.com/micode-underscore/episodes/s5e7-sommes-nous-a-laube-dun-effondrement-des-ia).]]>
            </summary>
            <updated>2025-08-28T21:01:56+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2091</id>
            <title type="text"><![CDATA[Crawl4AI]]></title>
            <link rel="alternate" href="https://crawl4ai.com/mkdocs/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2091"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open-Source LLM-Friendly Web Crawler &amp;amp; Scraper.

Crawl4AI delivers blazing-fast, AI-ready web crawling tailored for large language models, AI agents, and data pipelines. Fully open source, flexible, and built for real-time performance, Crawl4AI empowers developers with unmatched speed, precision, and deployment ease.

- [Crawl4AI @ GitHub](https://github.com/unclecode/crawl4ai).]]>
            </summary>
            <updated>2025-08-28T21:44:30+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2937</id>
            <title type="text"><![CDATA[Crawlee]]></title>
            <link rel="alternate" href="https://crawlee.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2937"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Build reliable crawlers. Fast.

A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. 

- [Crawlee @ GitHub](https://github.com/apify/crawlee).]]>
            </summary>
            <updated>2025-08-29T00:07:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2938</id>
            <title type="text"><![CDATA[Crawlee for Python]]></title>
            <link rel="alternate" href="https://crawlee.dev/python/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2938"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Build your Python web crawlers using Crawlee.
It helps you build reliable Python web crawlers. Fast.

 Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. 

- [Crawlee for Python @ GitHub](https://github.com/apify/crawlee-python).]]>
            </summary>
            <updated>2025-08-29T00:07:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3666</id>
            <title type="text"><![CDATA[Firecrawl]]></title>
            <link rel="alternate" href="https://www.firecrawl.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3666"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Turn websites into LLM-ready data.

Power your AI apps with clean data crawled from any website. It&amp;#039;s also open-source.
 🔥 Turn entire websites into LLM-ready markdown or structured data. Scrape, crawl and extract with a single API. 

- [Firecrawl @ GitHub](https://github.com/mendableai/firecrawl).

Related contents:

- [Firecrawl @ GitHub](https://github.com/firecrawl/firecrawl).
- [Firecrawl Observer @ GitHub](https://github.com/firecrawl/firecrawl-observer).

Related contents:

- [🚨 Someone built a tool that turns any website into clean data your AI can actually use @ Nav Toor&amp;#039;s X](https://nitter.net/heynavtoor/status/2031626457110425760).
- [Hermes Agent : veille technique auto-hébergée avec Matrix, FreshRSS et Firecrawl @ Cryptolab :fr:](https://cryptolab.re/posts/2026/hermes-agent-framework-self-hosted/).]]>
            </summary>
            <updated>2026-06-23T05:55:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3997</id>
            <title type="text"><![CDATA[Fess]]></title>
            <link rel="alternate" href="https://fess.codelibs.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3997"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Fess is very powerful and easily deployable Enterprise Search Server. 

Fess is a very powerful and easily deployable Enterprise Search Server. You can quickly install and run Fess on any platform where you can run the Java Runtime Environment. Fess is provided under the Apache License 2.0.

Fess is based on OpenSearch/Elasticsearch, but knowledge/experience about OpenSearch/Elasticsearch is not required. Fess provides an easy to use Administration GUI to configure the system via your browser. Fess also contains a Crawler, which can crawl documents on a web server, file system, or Data Store (such as a CSV or database). Many file formats are supported including (but not limited to): Microsoft Office, PDF, and zip.

- [Fess @ GitHub](https://github.com/codelibs/fess).
- [Docker for Fess @ GitHub](https://github.com/codelibs/docker-fess/).]]>
            </summary>
            <updated>2025-08-29T03:02:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6530</id>
            <title type="text"><![CDATA[Katana]]></title>
            <link rel="alternate" href="https://github.com/projectdiscovery/katana" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6530"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A next-generation crawling and spidering framework]]>
            </summary>
            <updated>2025-08-29T10:05:34+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7010</id>
            <title type="text"><![CDATA[Greenflare SEO Web Crawler]]></title>
            <link rel="alternate" href="https://greenflare.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7010"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The Open Source SEO Crawler]]>
            </summary>
            <updated>2025-08-29T11:26:20+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7061</id>
            <title type="text"><![CDATA[Screaming Frog SEO Spider Website Crawler]]></title>
            <link rel="alternate" href="https://www.screamingfrog.co.uk/seo-spider/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7061"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The industry leading website crawler for Windows, macOS and Ubuntu, trusted by thousands of SEOs and agencies worldwide for technical SEO site audits.]]>
            </summary>
            <updated>2025-08-29T11:34:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8764</id>
            <title type="text"><![CDATA[Scrapy]]></title>
            <link rel="alternate" href="http://scrapy.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8764"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[An open source web scraping framework for Python.

Scrapy is a fast high-level screen scraping and web crawling framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

- [Scrapy @ GitHub](https://github.com/scrapy/scrapy)]]>
            </summary>
            <updated>2025-08-29T16:18:56+00:00</updated>
        </entry>
    </feed>
