<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>scraping</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/171/feed"/>
    <updated>2026-09-17T19:03:33+00:00</updated>
    <id>https://links.biapy.com/guest/tags/171/feed</id>
            <entry>
            <id>https://links.biapy.com/links/13891</id>
            <title type="text"><![CDATA[Figranium]]></title>
            <link rel="alternate" href="https://figranium.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13891"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Build complex browser workflows visually.
Stack blocks visually to build complex browser workflows and execute them via API.

Figranium is an open-source visual browser automation platform with a block-based editor, built-in scheduling, proxy rotation, and a REST API to run everything.

- [Figranium @ GitHub](https://github.com/figranium/figranium).]]>
            </summary>
            <updated>2026-09-14T05:59:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13862</id>
            <title type="text"><![CDATA[Kitesurf]]></title>
            <link rel="alternate" href="https://kitesurf.cloudflare.app/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13862"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[stateless browser running entirely on Workers.

 Kitesurf is Cloudflare’s new stateless, highly scalable and cost-effective web browser that runs entirely on top of Workers and was designed specifically for the Agentic Cloud. Use this playground to explore Kitesurf capabilities. We inject Chrome DevTools in the UI, so you can inspect expanded DOM elements, read console messages, and watch network activity while Kitesurf renders pages. The Memory panel reports the WebAssembly footprint of each isolate, including frames, so you can gain a clear understanding of the resources each page is consuming. 

Related contents:

- [Introducing Kitesurf: The agent-first browser that runs in V8 isolates on Cloudflare Workers @ Cloudflare blog](https://blog.cloudflare.com/kitesurf/).]]>
            </summary>
            <updated>2026-09-10T06:30:55+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13861</id>
            <title type="text"><![CDATA[Obscura]]></title>
            <link rel="alternate" href="https://obscura.sh/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13861"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Give every agent its own browser.

The headless browser for AI agents and web scraping.
Obscura is a headless browser engine written in Rust, built for web scraping and AI agent automation. It runs real JavaScript via V8, supports the Chrome DevTools Protocol, and acts as a drop-in replacement for headless Chrome with Puppeteer and Playwright.

- [Obscura @ GitHub](https://github.com/h4ckf0r0day/obscura).]]>
            </summary>
            <updated>2026-09-10T06:28:47+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13751</id>
            <title type="text"><![CDATA[Accept: text/markdown]]></title>
            <link rel="alternate" href="https://acceptmarkdown.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13751"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Serve Markdown to AI Agents with  `Accept: text/markdown`

Your site already has the content. Serving a Markdown variant lets agent AI clients skip nav/scripts/layout markup and read the content directly.]]>
            </summary>
            <updated>2026-09-02T16:50:52+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13688</id>
            <title type="text"><![CDATA[BetterWright]]></title>
            <link rel="alternate" href="https://betterwright.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13688"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Token-Efficient Browser for AI Agents.

A persistent, policy-guarded Playwright browser for AI agents — network policy, encrypted credential vault, proof screenshots, and CAPTCHA solving. 

- [BetterWright @ GitHub](https://github.com/BetterWright/betterwright).]]>
            </summary>
            <updated>2026-08-20T11:28:48+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13662</id>
            <title type="text"><![CDATA[Moli Browser]]></title>
            <link rel="alternate" href="https://browser.lexmount.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13662"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Moli is a production-ready headless browser for AI agents. Its on-demand layout and rendering design combines a complete browser runtime with a lightweight resource footprint.

- [Moli Browser @ GitHub](https://github.com/lexmount/moli).]]>
            </summary>
            <updated>2026-08-17T08:40:32+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13336</id>
            <title type="text"><![CDATA[wigolo]]></title>
            <link rel="alternate" href="https://knockoutez.github.io/wigolo/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13336"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The web, wired into your local agent.

wigolo is a local-first server that hands any AI agent the whole web — search, fetch, crawl, extract, cache, and research. In your editor over MCP, in your framework through an SDK, or in your self-hosted stack over REST. No API keys. No cloud. No metered bill.

- [wigolo @ GitHub](https://github.com/KnockOutEZ/wigolo).]]>
            </summary>
            <updated>2026-07-20T12:03:45+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13302</id>
            <title type="text"><![CDATA[👁️ Agent Reach :cn:]]></title>
            <link rel="alternate" href="https://github.com/Panniantong/Agent-Reach/tree/main" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13302"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Give your AI agent eyes to see the entire internet. Read &amp;amp; search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.]]>
            </summary>
            <updated>2026-07-15T08:40:52+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13280</id>
            <title type="text"><![CDATA[mimic]]></title>
            <link rel="alternate" href="https://github.com/littledivy/mimic" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13280"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Intercept any app, then call it from Python like a library.]]>
            </summary>
            <updated>2026-07-15T06:49:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13114</id>
            <title type="text"><![CDATA[kage]]></title>
            <link rel="alternate" href="https://kage.tamnd.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13114"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A website, frozen as a shadow.
 Shadow any website for offline viewing, with the JavaScript stripped out .

kage renders every page in headless Chrome, snapshots the final DOM, removes every script and event handler, and downloads and rewrites the CSS, images, and fonts. The result looks like the live site but runs no code: a plain folder of .html files you can open straight from disk.

- [kage @ GitHub](https://github.com/tamnd/kage).]]>
            </summary>
            <updated>2026-06-25T15:40:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13010</id>
            <title type="text"><![CDATA[.MD this page]]></title>
            <link rel="alternate" href="https://github.com/Ademking/MD-This-Page" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13010"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Convert any web page to clean, readable Markdown with just one click. 

Turn any webpage into clean, LLM-ready Markdown in one click.
Strip the browser, keep the structure, then copy or download instantly.]]>
            </summary>
            <updated>2026-06-12T11:46:58+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/13001</id>
            <title type="text"><![CDATA[SerpApi]]></title>
            <link rel="alternate" href="https://serpapi.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/13001"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Google Search API.

Scrape Google and other search engines from our fast, easy, and complete API.

- [SerpApi @ GitHub](https://github.com/serpapi/).

Related contents:

- [Web Scraping for Beginners – Extract Data with an API @ freeCodeCamp.org&amp;#039;s YouTube](https://www.youtube.com/watch?v=j6hnjNhx_MM).]]>
            </summary>
            <updated>2026-06-12T07:12:24+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12984</id>
            <title type="text"><![CDATA[Reader]]></title>
            <link rel="alternate" href="https://reader.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12984"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AI Web Infrastructure Platform.

 Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents. 

- [Reader @ GitHub](https://github.com/vakra-dev/reader).

Related contents:

- [Top 19 des alternatives open-source aux SaaS en 2026 @ Camille Roux :fr:](https://www.camilleroux.com/alternatives-open-source-saas-2026-v2/).]]>
            </summary>
            <updated>2026-06-10T12:46:46+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12867</id>
            <title type="text"><![CDATA[Crawl4AI]]></title>
            <link rel="alternate" href="https://docs.crawl4ai.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12867"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open-source LLM Friendly Web Crawler &amp;amp; Scraper.

Crawl4AI turns the web into clean, LLM ready Markdown for RAG, agents, and data pipelines. Fast, controllable, battle tested by a 50k+ star community.

- [Crawl4AI @ GitHub](https://github.com/unclecode/crawl4ai).]]>
            </summary>
            <updated>2026-08-07T06:12:30+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12711</id>
            <title type="text"><![CDATA[curl.md]]></title>
            <link rel="alternate" href="https://curl.md/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12711"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[URL to markdown for agents.

Turn websites into optimized, low token output to supercharge your context. Works with every agent.

- [curl.md @ GitHub](https://github.com/wevm/curl.md).]]>
            </summary>
            <updated>2026-05-14T14:54:11+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12678</id>
            <title type="text"><![CDATA[Obscura]]></title>
            <link rel="alternate" href="https://github.com/h4ckf0r0day/obscura" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12678"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The headless browser for AI agents and web scraping.
 Lightweight, stealthy, and built in Rust. 

Obscura is a headless browser engine written in Rust, built for web scraping and AI agent automation. It runs real JavaScript via V8, supports the Chrome DevTools Protocol, and acts as a drop-in replacement for headless Chrome with Puppeteer and Playwright.]]>
            </summary>
            <updated>2026-05-14T08:05:28+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12618</id>
            <title type="text"><![CDATA[Is Your Site Agent-Ready?]]></title>
            <link rel="alternate" href="https://isitagentready.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12618"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Scan your website to see how ready it is for AI agents. We check multiple emerging standards — from robots.txt and Markdown negotiation to MCP, OAuth, Agent Skills and agentic commerce.

Related contents:

- [#140: News septembre 2026 RC2, Tailwind rejoint Shopify, des tokens gratuits et Homebrew UI @ Double Slash :fr:](https://double-slash.dev/podcasts/news-sept26-rc2/).]]>
            </summary>
            <updated>2026-09-17T06:42:03+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12482</id>
            <title type="text"><![CDATA[skim]]></title>
            <link rel="alternate" href="https://github.com/noblepayne/skim" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12482"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Turns any HTML page into clean markdown. 

Mozilla Readability algorithm + hickory-to-markdown converter for Babashka.

Turns any HTML page into clean markdown by parsing it with jsoup, scoring content using the Mozilla Readability algorithm, extracting the main article body, and converting the hickory tree to markdown.

Related contents:

- [Episode 661: Sink Your Claws In @ Linux Unplugged](https://linuxunplugged.com/661).]]>
            </summary>
            <updated>2026-04-09T06:24:14+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12281</id>
            <title type="text"><![CDATA[Defuddle]]></title>
            <link rel="alternate" href="https://defuddle.md/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12281"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Get the main content of any page as Markdown.

Defuddle extracts the main content from web pages. It cleans up web pages by removing clutter like comments, sidebars, headers, footers, and other non-essential elements, leaving only the primary content.

- [Defuddle @ GitHub](https://github.com/kepano/defuddle).

Related contents:

- [1,65 million de vues en 30 jours : comment j&amp;#039;ai automatisé la diffusion de ma veille techno @ Camille Roux :fr:](https://www.camilleroux.com/1-65-million-de-vues-en-30-jours-comment-jai-automatise-la-diffusion-de-ma-veille-techno/).]]>
            </summary>
            <updated>2026-03-24T13:56:42+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12272</id>
            <title type="text"><![CDATA[Scrapy]]></title>
            <link rel="alternate" href="https://www.scrapy.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12272"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Scrapy, a fast high-level web crawling &amp;amp; scraping framework for Python.

- [Scrapy @ GitHub](https://github.com/scrapy/scrapy).]]>
            </summary>
            <updated>2026-03-24T08:31:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12227</id>
            <title type="text"><![CDATA[CloakBrowser]]></title>
            <link rel="alternate" href="https://cloakbrowser.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12227"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Stealth Chromium for Browser Automation.
Stealth Chromium that passes every bot detection test.

Not a patched config. Not a JS injection. A real Chromium binary with fingerprints modified at the C++ source level. Antibot systems score it as a normal browser — because it is a normal browser.

- [CloakBrowser @ GitHub](https://github.com/CloakHQ/CloakBrowser).]]>
            </summary>
            <updated>2026-03-20T14:11:36+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/12074</id>
            <title type="text"><![CDATA[haphash]]></title>
            <link rel="alternate" href="https://github.com/dgl/haphash" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12074"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Anti-scraper challenge for haproxy to stop naughty AI bots. 

This is a simple anti-scraper solution for haproxy, using a similar &amp;quot;hashcash&amp;quot; challenge as anubis uses. The goal is to be as simple as possible, so this can be implemented alongside other haproxy rules to control traffic.]]>
            </summary>
            <updated>2026-03-10T08:26:08+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11945</id>
            <title type="text"><![CDATA[Scrapling]]></title>
            <link rel="alternate" href="https://scrapling.readthedocs.io/en/latest/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11945"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! 

Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box. And its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume and automatic proxy rotation — all in a few lines of Python. One library, zero compromises.

- [Scrapling @ GitHub](https://github.com/D4Vinci/Scrapling).

Related contents:

- [Scrapling - Le scraper Python qui se répare tout seul @ Korben :fr:](https://korben.info/scrapling-scraper-python-auto-repare.html).]]>
            </summary>
            <updated>2026-04-28T09:04:16+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11874</id>
            <title type="text"><![CDATA[Dembrandt]]></title>
            <link rel="alternate" href="https://www.dembrandt.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11874"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Source CLI to Extract Design Systems &amp;amp; Tokens.

Extract any website’s design system into design tokens in a few seconds: logo, colors, typography, borders, and more. One command.

Extract design tokens from any website with one command. CSS to design system converter with W3C design tokens export.

- [Dembrandt @ GitHub](https://github.com/dembrandt/dembrandt).]]>
            </summary>
            <updated>2026-02-20T07:34:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11583</id>
            <title type="text"><![CDATA[Doppelganger]]></title>
            <link rel="alternate" href="https://doppelgangerdev.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11583"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Browser automation for everyone.

Doppelganger is a self-hosted, free browser automation and scraping platform built on Playwright, designed to handle everything from simple scraping to complex, human-like browser interactions. It provides a visual task editor, a structured JSON task format, and advanced execution modes that go far beyond traditional scrapers. 

- [Doppelganger @ GitHub](https://github.com/mnemosyne-artificial-intelligence/doppelganger).]]>
            </summary>
            <updated>2026-01-23T13:46:54+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11557</id>
            <title type="text"><![CDATA[wxpath]]></title>
            <link rel="alternate" href="https://github.com/rodricios/wxpath" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11557"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[declarative web crawling with XPath.

wxpath is a declarative web crawler where traversal is expressed directly in XPath. Instead of writing imperative crawl loops, wxpath lets you describe what to follow and what to extract in a single expression. wxpath executes that expression concurrently, breadth-first-ish, and streams results as they are discovered.]]>
            </summary>
            <updated>2026-01-21T13:06:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11387</id>
            <title type="text"><![CDATA[shot-scraper]]></title>
            <link rel="alternate" href="https://shot-scraper.datasette.io/en/stable/index.html" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11387"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A command-line utility for taking automated screenshots of websites.

- [shot-scraper @ GitHub](https://github.com/simonw/shot-scraper).

Related contents:

- [How Rob Pike got spammed with an AI slop “act of kindness” @ Simon Willison’s Weblog](https://simonwillison.net/2025/Dec/26/slop-acts-of-kindness/).]]>
            </summary>
            <updated>2026-01-06T07:07:44+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11225</id>
            <title type="text"><![CDATA[iocaine]]></title>
            <link rel="alternate" href="https://iocaine.madhouse-project.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11225"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[the deadliest poison known to AI.

This software is not made for making the Crawlers go away. It is an aggressive defense mechanism that tries its best to take the blunt of the assault, serve them garbage, and keep them off of upstream resources. Even though a lot of work went into making iocaine efficient, and nigh invisible for the legit visitor, it is an aggressive defender nevertheless, and will require a few resources - a whole lot less than if you’d let the Crawlers run rampant, though.

- [iocaine &amp;#039;s git](https://git.madhouse-project.org/iocaine/iocaine).

Related contents:

- [Guarding My Git Forge Against AI Scrapers @ VulpineCitrus](https://vulpinecitrus.info/blog/guarding-git-forge-ai-scrapers/).]]>
            </summary>
            <updated>2025-12-12T13:43:40+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/11027</id>
            <title type="text"><![CDATA[Markdown Printer]]></title>
            <link rel="alternate" href="https://github.com/levz0r/markdown-printer" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11027"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Save web pages as Markdown files with preserved formatting. Perfect for documentation, articles, and note-taking. 

Save web pages as Markdown files with preserved formatting. Zero setup required - just install and start saving!]]>
            </summary>
            <updated>2025-11-20T16:10:50+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10818</id>
            <title type="text"><![CDATA[Herr Bischoff&amp;#039;s Bot Database]]></title>
            <link rel="alternate" href="https://badbot.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10818"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Please find below a manually curated and researched list of users agents I came across. It&amp;#039;s impressive to see how many of the bots active today flat out do not respect robots.txt settings — or claim to do it but ignore them. This list is updated regularly, whenever I spot new user agents and look into their behavior. There is no JavaScript, here no fancy search.

Related contents:

- [Comment protéger vos serveurs et lutter efficacement contre les crawlers d’IA @ Bearstech :fr:](https://bearstech.com/societe/blog/comment-proteger-vos-serveurs-et-lutter-efficacement-contre-les-crawlers-dia).]]>
            </summary>
            <updated>2025-10-30T06:48:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10495</id>
            <title type="text"><![CDATA[ai.robots.txt]]></title>
            <link rel="alternate" href="https://github.com/ai-robots-txt/ai.robots.txt" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10495"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A list of AI agents and robots to block. 

This list contains AI-related crawlers of all types, regardless of purpose. We encourage you to contribute to and implement this list on your own site. See information about the listed crawlers and the FAQ.

Related contents:

- [\#119: Les news sur le développement web et l&amp;#039;IA pour septembre 2025 RC2 @ Double Slash :fr:](https://double-slash.dev/podcasts/news-sept25-rc2/).
- [Comment bloquer les crawlers IA qui pillent votre site sans vous demander la permission ? @ Korben :fr:](https://korben.info/bloquer-crawlers-ia-robots-txt-htaccess-nginx.html).]]>
            </summary>
            <updated>2025-12-16T10:41:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10251</id>
            <title type="text"><![CDATA[The /llms.txt file]]></title>
            <link rel="alternate" href="https://llmstxt.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10251"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A proposal to standardise on using an /llms.txt file to provide information to help LLMs use a website at inference time.

We propose adding a /llms.txt markdown file to websites to provide LLM-friendly content. This file offers brief background information, guidance, and links to detailed markdown files.

llms.txt markdown is human and LLM readable, but is also in a precise format allowing fixed processing methods (i.e. classical programming techniques such as parsers and regex).

- [llms.txt @ GitHub](https://github.com/answerdotai/llms-txt).
- [MCP LLMS-TXT Documentation Server @ GitHub](https://github.com/langchain-ai/mcpdoc).

Related contents:

- [Firecrawl LLMs.txt Generator @ GitHub](https://github.com/firecrawl/create-llmstxt-py).
- [A proposal for inline LLM instructions in HTML based on llms.txt @ Vercel](https://vercel.com/blog/a-proposal-for-inline-llm-instructions-in-html).
- [\#119: Les news sur le développement web et l&amp;#039;IA pour septembre 2025 RC2 @ Double Slash :fr:](https://double-slash.dev/podcasts/news-sept25-rc2/).
- [A fun trick for getting discovered by LLMs and AI tools @ Cassidy Williams](https://cassidoo.co/post/ai-llm-discoverability/).]]>
            </summary>
            <updated>2026-02-02T10:24:56+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10217</id>
            <title type="text"><![CDATA[mdream]]></title>
            <link rel="alternate" href="https://github.com/harlan-zw/mdream" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10217"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[☁️ Convert any site to clean markdown &amp;amp; llms.txt. Boost your site&amp;#039;s AI discoverability or generate LLM context for a project you&amp;#039;re working with. 

Mdream core is a highly optimized primitive for producing Markdown from HTML that is optimized for LLMs.]]>
            </summary>
            <updated>2025-09-15T14:08:37+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10197</id>
            <title type="text"><![CDATA[dirsearch]]></title>
            <link rel="alternate" href="https://github.com/maurosoria/dirsearch" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10197"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Web path scanner.
An advanced web path brute-forcer.

Related contents:

- [DirSearch - Un scanner de chemins web @ Korben :fr:](https://korben.info/2025-09-12-dirsearch-scanner-web-paths.html).]]>
            </summary>
            <updated>2025-09-15T09:23:10+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10032</id>
            <title type="text"><![CDATA[Byparr]]></title>
            <link rel="alternate" href="https://github.com/ThePhaseless/Byparr" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10032"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Get your valid antibot cookies yourself!

Built with seleniumbase and FastAPI, this project aims to mimic FlareSolverr&amp;#039;s API and functionality of providing you with http cookies and headers for websites protected with anti-bot protections.

Related contents:

- [Byparr bypasses Flaresolverr @ ElfHosted](https://store.elfhosted.com/blog/2025/04/16/byparr-bypasses-flaresolverr/).]]>
            </summary>
            <updated>2025-09-05T19:31:57+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/59</id>
            <title type="text"><![CDATA[🤖 Botasaurus Framework 🤖]]></title>
            <link rel="alternate" href="https://www.omkar.cloud/botasaurus/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/59"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Botasaurus is a Swiss Army knife 🔪 for web scraping and browser automation 🤖 that helps you create bots fast. ⚡️

Botasaurus is an all-in-one web scraping framework that enables you to build awesome scrapers in less time, with less code, and with more fun.

- [🤖 Botasaurus 🤖 @ GitHub](https://github.com/omkarcloud/botasaurus).

Related contents:

- [Botasaurus - Le scraper qui rend Cloudflare aussi facile à contourner qu&amp;#039;un CAPTCHA de 2005 @ Korben :fr:](https://korben.info/botasaurus-framework-python-rend-cloudflare-aussi.html).]]>
            </summary>
            <updated>2025-09-04T08:40:18+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/193</id>
            <title type="text"><![CDATA[A Vocabulary For Expressing AI Usage Preferences]]></title>
            <link rel="alternate" href="https://ietf-wg-aipref.github.io/drafts/draft-ietf-aipref-vocab.html?cf_target_id=_blank" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/193"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[This document proposes a standardized vocabulary for expressing preferences related to how digital assets are used by automated processing systems. This vocabulary allows for the creation of structured declarations about restrictions or permissions for use of digital assets by such systems.

Related contents:

- [Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives @ Cloudflare Blog](https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/).]]>
            </summary>
            <updated>2025-09-18T15:26:23+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/247</id>
            <title type="text"><![CDATA[Meka Agent]]></title>
            <link rel="alternate" href="https://withmeka.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/247"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[state of the art browsing agent (WebArena 72.7%).

Meka Agent is an open-source, autonomous computer-using agent that delivers state-of-the-art browsing capabilities. The agent works and acts in the same way humans do, by purely using vision as its eyes and acting within a full computer context.

It is designed as a simple, extensible, and customizable framework, allowing flexibility in the choice of models, tools, and infrastructure providers.

- [Meka Agent @ GitHub](https://github.com/trymeka/agent).]]>
            </summary>
            <updated>2025-12-09T11:38:26+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/339</id>
            <title type="text"><![CDATA[Fireplexity]]></title>
            <link rel="alternate" href="https://tools.firecrawl.dev/fireplexity" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/339"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AI-Powered Web Scraping &amp;amp; Data Enrichment.
AI-powered web search with instant results and follow-up questions.

 🔥 Blazing-fast AI search engine with real-time citations, streaming responses, and live data powered by Firecrawl 

- [Fireplexity @ GitHub](https://github.com/mendableai/fireplexity).

Related contents:

- [\#116: Les news sur le développement web et l&amp;#039;IA pour juillet 2025 RC1@ Double Slash :fr:](https://double-slash.dev/podcasts/news-jul25/).]]>
            </summary>
            <updated>2026-01-21T13:12:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/352</id>
            <title type="text"><![CDATA[The Web Robots Pages]]></title>
            <link rel="alternate" href="https://www.robotstxt.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/352"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Web Robots (also known as Web Wanderers, Crawlers, or Spiders), are programs that traverse the Web automatically. Search engines such as Google use them to index the web content, spammers use them to scan for email addresses, and they have many other uses.

Related contents:

- [I was wrong about robots.txt @ Evgenii Pendragon](https://evgeniipendragon.com/posts/i-was-wrong-about-robots-txt/).
- [Fix Your robots.txt or Your Site Disappears from Google @ alanwsmith.com](https://www.alanwsmith.com/en/37/wa/jz/s1/).]]>
            </summary>
            <updated>2026-01-22T12:47:53+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/541</id>
            <title type="text"><![CDATA[Maxun Cloud]]></title>
            <link rel="alternate" href="https://www.maxun.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/541"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[No-Code Web Data Extraction Platform.
Turn Websites To APIs &amp;amp; Spreadsheets In Minutes.

Maxun lets you train a robot in 2 minutes and scrape the web on auto-pilot. Web data extraction doesn&amp;#039;t get easier than this! 

- [Maxun @ GitHub](https://github.com/getmaxun/maxun).]]>
            </summary>
            <updated>2025-08-28T17:27:58+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/593</id>
            <title type="text"><![CDATA[Scraperr]]></title>
            <link rel="alternate" href="https://scraperr-docs.pages.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/593"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A powerful self-hosted web scraping solution.
Scrape websites without writing a single line of code.

- [Scraperr @ GitHub](https://github.com/jaypyles/Scraperr).]]>
            </summary>
            <updated>2026-03-13T13:36:47+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/656</id>
            <title type="text"><![CDATA[Defuddle]]></title>
            <link rel="alternate" href="https://kepano.github.io/defuddle/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/656"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Defuddle extracts the main content from web pages. It cleans up web pages by removing clutter like comments, sidebars, headers, footers, and other non-essential elements, leaving only the primary content.

- [Defuddle @ GitHub](https://github.com/kepano/defuddle).]]>
            </summary>
            <updated>2025-08-28T17:47:08+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/799</id>
            <title type="text"><![CDATA[CleverBee]]></title>
            <link rel="alternate" href="https://cleverb.ee/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/799"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Transparent AI, Rooted in Research, Open to All.
The Open Source Deep Researcher Tool.
AI-Powered Online Data Information Synthesis Assistant.

CleverBee is a powerful Python-based research assistant agent using Large Language Models (LLMs) like Claude and Gemini, Playwright for web browsing, and Chainlit for an interactive UI. It performs research assistance by browsing the web, extracting content (HTML), cleaning it, and synthesizing findings based on user research topics.

- [CleverBee @ GitHub](https://github.com/SureScaleAI/cleverbee).]]>
            </summary>
            <updated>2025-08-28T18:10:29+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/864</id>
            <title type="text"><![CDATA[TeleGraphite]]></title>
            <link rel="alternate" href="https://github.com/hamodywe/telegram-scraper-TeleGraphite" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/864"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Telegram Scraper &amp;amp; JSON Exporter &amp;amp; telegram chanels scraper.

 A fast and reliable Telegram channel scraper that fetches posts and exports them to JSON.]]>
            </summary>
            <updated>2025-08-28T18:22:29+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1132</id>
            <title type="text"><![CDATA[PriceBuddy]]></title>
            <link rel="alternate" href="https://github.com/jez500/pricebuddy" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1132"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[PriceBuddy is an open source, self-hostable, web application that allows users to compare prices of products from different online retailers. Users can search for a product and view the prices of that product from different online retailers.]]>
            </summary>
            <updated>2025-08-28T19:04:50+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1218</id>
            <title type="text"><![CDATA[Anubis]]></title>
            <link rel="alternate" href="https://anubis.techaro.lol/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1218"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Anubis: self hostable scraper defense software.

Weighs the soul of incoming HTTP requests using proof-of-work to stop AI crawlers.

- [Anubis @ GitHub](https://github.com/TecharoHQ/anubis).

Related contents:

- [Block AI scrapers with Anubis @ Xe](https://xeiaso.net/blog/2025/anubis/).
- [Episode 146: When AI Attacks @ Self-Hosted](https://selfhosted.show/146).
- [The surreal joy of having an overprovisioned homelab @ Xe](https://xeiaso.net/talks/2025/surreal-joy-homelab/).
- [Open source devs are fighting AI crawlers with cleverness and vengeance @ TechCrunch](https://techcrunch.com/2025/03/27/open-source-devs-are-fighting-ai-crawlers-with-cleverness-and-vengeance/).
- [\[Anubis\] Utiliser la preuve de travail pour bloquer les robots @ Pofilo.fr :fr:](https://www.pofilo.fr/post/2025/04/14-mise-en-place-anubis/).
- [The Day Anubis Saved Our Websites From a DDoS Attack @ fabulous.systems](https://fabulous.systems/posts/2025/05/anubis-saved-our-websites-from-a-ddos-attack/).
- [Protéger tous ses sites avec Anubis @ Dryusdan.space 🚀](https://dryusdan.space/proteger-tous-ses-sites-avec-anubis).
- [A thought on JavaScript &amp;quot;proof of work&amp;quot; anti-scraper systems @ Wandering Thoughts](https://utcc.utoronto.ca/~cks/space/blog/web/JavaScriptScraperObstacles).
- [Anubis - Protégez votre site web contre les scrapers IA en moins de 15 minutes @ Korben :fr:](https://korben.info/anubis-protection-site-web-bots-ia.html).
- [Ask HN: How to stop an AWS bot sending 2B requests/month? @ Hacker News](https://news.ycombinator.com/item?id=45613567).
- [Comment protéger vos serveurs et lutter efficacement contre les crawlers d’IA @ Bearstech :fr:](https://bearstech.com/societe/blog/comment-proteger-vos-serveurs-et-lutter-efficacement-contre-les-crawlers-dia).]]>
            </summary>
            <updated>2025-10-30T06:49:52+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1223</id>
            <title type="text"><![CDATA[Fetcher MCP]]></title>
            <link rel="alternate" href="https://github.com/jae-jae/fetcher-mcp" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1223"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[MCP server for fetch web page content using Playwright headless browser.]]>
            </summary>
            <updated>2025-08-28T19:19:58+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1803</id>
            <title type="text"><![CDATA[Lightpanda Browser]]></title>
            <link rel="alternate" href="https://lightpanda.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1803"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The headless browser. The future of browser automation.

Combining the flexibility of a Headless Chrome API with an innovative tailor-made browser for unmatched performance and efficiency in browser automation.

- [Lightpanda Browser @ GitHub](https://github.com/lightpanda-io/browser).

Related contents:

- [Lightpanda - Le navigateur rapide pour l&amp;#039;automatisation web @ Korben :fr:](https://korben.info/lightpanda-navigateur-automatisation-web.html).]]>
            </summary>
            <updated>2026-03-17T07:47:20+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1827</id>
            <title type="text"><![CDATA[Nepenthes]]></title>
            <link rel="alternate" href="https://zadzmo.org/code/nepenthes/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1827"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[This is a tarpit intended to catch web crawlers. Specifically, it&amp;#039;s targetting crawlers that scrape data for LLM&amp;#039;s - but really, like the plants it is named after, it&amp;#039;ll eat just about anything that finds it&amp;#039;s way inside.

It works by generating an endless sequences of pages, each of which with dozens of links, that simply go back into a the tarpit. Pages are randomly generated, but in a deterministic way, causing them to appear to be flat files that never change. Intentional delay is added to prevent crawlers from bogging down your server, in addition to wasting their time. Lastly, optional Markov-babble can be added to the pages, to give the crawlers something to scrape up and train their LLMs on, hopefully accelerating model collapse.

Related contents:

- [Nepenthes - Piégez les crawlers web malveillants @ Korben :fr:](https://korben.info/nepenthes-piege-crawlers-web-malveillants.html).
- [Open source devs are fighting AI crawlers with cleverness and vengeance @ TechCrunch](https://techcrunch.com/2025/03/27/open-source-devs-are-fighting-ai-crawlers-with-cleverness-and-vengeance/).
- [Ask HN: How to stop an AWS bot sending 2B requests/month? @ Hacker News](https://news.ycombinator.com/item?id=45613567).]]>
            </summary>
            <updated>2025-10-20T06:35:24+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1835</id>
            <title type="text"><![CDATA[Common Crawl]]></title>
            <link rel="alternate" href="https://commoncrawl.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1835"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Repository of Web Crawl Data.

Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.

Related contents:

- [S5E7 - Sommes-nous à l&amp;#039;aube d&amp;#039;un effondrement des IA ? @ Underscore_&amp;#039;s acast :fr:](https://shows.acast.com/micode-underscore/episodes/s5e7-sommes-nous-a-laube-dun-effondrement-des-ia).]]>
            </summary>
            <updated>2025-08-28T21:01:56+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2036</id>
            <title type="text"><![CDATA[DrissionPage官网 :cn:]]></title>
            <link rel="alternate" href="https://drissionpage.cn/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2036"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Python based web automation tool. Powerful and elegant. 

DrissionPage is a Python-based web automation tool.
It can control the browser, send and receive packets, and combine the two.
You can balance the convenience of browser automation with the efficiency of requests.
It is powerful, built-in countless user-friendly design and convenient features.
Its syntax is simple and elegant, the code is small, and it is friendly to beginners.

- [DrissionPage @ GitHub](https://github.com/g1879/DrissionPage).]]>
            </summary>
            <updated>2025-08-28T21:36:19+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2091</id>
            <title type="text"><![CDATA[Crawl4AI]]></title>
            <link rel="alternate" href="https://crawl4ai.com/mkdocs/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2091"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open-Source LLM-Friendly Web Crawler &amp;amp; Scraper.

Crawl4AI delivers blazing-fast, AI-ready web crawling tailored for large language models, AI agents, and data pipelines. Fully open source, flexible, and built for real-time performance, Crawl4AI empowers developers with unmatched speed, precision, and deployment ease.

- [Crawl4AI @ GitHub](https://github.com/unclecode/crawl4ai).]]>
            </summary>
            <updated>2025-08-28T21:44:30+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2132</id>
            <title type="text"><![CDATA[AIL-Framework]]></title>
            <link rel="alternate" href="https://github.com/supdevinci/ail-framework-docker" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2132"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AIL-Framework is a powerful open-source project designed for online data analysis and web crawling, tailored for cybersecurity researchers and analysts.

Related contents:

- [1 Tools en 5 commandes @ Laurent Biagotti&amp;#039;s LinkedIn :fr:](https://www.linkedin.com/posts/laurent-biagiotti-19779284_1-%F0%9D%97%A7%F0%9D%97%BC%F0%9D%97%BC%F0%9D%97%B9%F0%9D%98%80-en-5-%F0%9D%97%96%F0%9D%97%BC%F0%9D%97%BA%F0%9D%97%BA%F0%9D%97%AE%F0%9D%97%BB%F0%9D%97%B1%F0%9D%97%B2%F0%9D%98%80-et-activity-7281937762511929344-MiOX/).]]>
            </summary>
            <updated>2025-08-28T21:52:26+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2473</id>
            <title type="text"><![CDATA[DataFuel]]></title>
            <link rel="alternate" href="https://www.datafuel.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2473"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Turn websites into LLM - ready   data.

DataFuel API scrapes entire websites and knowledge bases in a single query. Get clean, markdown-structured web data instantly for your RAG systems and AI models. No complex scraping code needed.]]>
            </summary>
            <updated>2025-08-28T22:48:59+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2497</id>
            <title type="text"><![CDATA[Darkdump]]></title>
            <link rel="alternate" href="https://github.com/josh0xA/darkdump" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2497"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open Source Intelligence Interface for Deep Web Scraping.

Darkdump is a OSINT interface for carrying out deep web investgations written in python in which it allows users to enter a search query in which darkdump provides the ability to scrape .onion sites relating to that query to try to extract emails, metadata, keywords, images, social media etc. Darkdump retrieves sites via Ahmia.fi and scrapes those .onion addresses when connected via the tor network.

Related contents:

- [darkdump: Open Source Intelligence Interface for Deep Web Scraping @ Dark Web Informer](https://darkwebinformer.com/darkdump-open-source-intelligence-interface-for-deep-web-scraping/).
- [Darkdump - L&amp;#039;outil OSINT qui fouille le dark web pour vous @ Korben :fr:](https://korben.info/darkdump-outil-osint-fouille-dark-web.html).]]>
            </summary>
            <updated>2025-08-28T22:53:00+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2717</id>
            <title type="text"><![CDATA[Discount Bandit]]></title>
            <link rel="alternate" href="https://discount-bandit.cybrarist.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2717"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Self Hosted product tracker for Amazon, Walmart And many more.

 Track products pricing across multi ecommerce stores such as amazon,ebay,walmart, target and many more. 

- [Discount Bandit @ GitHub](https://github.com/Cybrarist/Discount-Bandit).
- [Episode 590: Self-Host Before You&amp;#039;re Toast @ Linux Unplugged](https://linuxunplugged.com/590).]]>
            </summary>
            <updated>2025-08-28T23:29:20+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2744</id>
            <title type="text"><![CDATA[simple-cloudflare-solver]]></title>
            <link rel="alternate" href="https://github.com/nlevee/simple-cloudflare-solver" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2744"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[simple-cloudflare-solver is an API to bypass Cloudflare&amp;#039;s protection system. It can be used as a gateway by applications like Jackett and Prowlarr to access protected resources.]]>
            </summary>
            <updated>2025-08-28T23:33:25+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2801</id>
            <title type="text"><![CDATA[GitDorker]]></title>
            <link rel="alternate" href="https://github.com/obheda12/GitDorker" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2801"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A Python program to scrape secrets from GitHub through usage of a large repository of dorks. 

GitDorker is a tool that utilizes the GitHub Search API and an extensive list of GitHub dorks that I&amp;#039;ve compiled from various sources to provide an overview of sensitive information stored on github given a search query.

The Primary purpose of GitDorker is to provide the user with a clean and tailored attack surface to begin harvesting sensitive information on GitHub. GitDorker can be used with additional tools such as GitRob or Trufflehog on interesting repos or users discovered from GitDorker to produce best results.]]>
            </summary>
            <updated>2025-08-28T23:43:24+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2883</id>
            <title type="text"><![CDATA[Scrapling]]></title>
            <link rel="alternate" href="https://github.com/D4Vinci/Scrapling" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2883"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Undetectable, Lightning-Fast, and Adaptive Web Scraping for Python.

Dealing with failing web scrapers due to anti-bot protections or website changes? Meet Scrapling.

Scrapling is a high-performance, intelligent web scraping library for Python that automatically adapts to website changes while significantly outperforming popular alternatives. For both beginners and experts, Scrapling provides powerful features while maintaining simplicity.]]>
            </summary>
            <updated>2025-08-28T23:57:35+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2937</id>
            <title type="text"><![CDATA[Crawlee]]></title>
            <link rel="alternate" href="https://crawlee.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2937"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Build reliable crawlers. Fast.

A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation. 

- [Crawlee @ GitHub](https://github.com/apify/crawlee).]]>
            </summary>
            <updated>2025-08-29T00:07:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/2938</id>
            <title type="text"><![CDATA[Crawlee for Python]]></title>
            <link rel="alternate" href="https://crawlee.dev/python/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/2938"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Build your Python web crawlers using Crawlee.
It helps you build reliable Python web crawlers. Fast.

 Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation. 

- [Crawlee for Python @ GitHub](https://github.com/apify/crawlee-python).]]>
            </summary>
            <updated>2025-08-29T00:07:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3002</id>
            <title type="text"><![CDATA[Maxun]]></title>
            <link rel="alternate" href="https://maxun-website.vercel.app/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3002"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Open-Source No-Code Web Data Extraction Platform.

Build custom robots to automate data scraping.

- [Maxun @ GitHub](https://github.com/getmaxun/maxun).]]>
            </summary>
            <updated>2025-08-29T00:16:22+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3173</id>
            <title type="text"><![CDATA[Scrappey]]></title>
            <link rel="alternate" href="https://scrappey.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3173"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Web Scraping API.
Tired of getting blocked while Scraping the web?

Our simple-to-use API makes it easy. Rotating proxies, Anti-Bot technology and headless browsers to CAPTCHAs. It&amp;#039;s never been this easy.]]>
            </summary>
            <updated>2025-08-29T00:44:55+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3174</id>
            <title type="text"><![CDATA[Cloudflare Turnstile Page &amp;amp; Captcha Bypass for Scraping]]></title>
            <link rel="alternate" href="https://github.com/sarperavci/CloudflareBypassForScraping" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3174"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A cloudflare verification bypass script for webscraping.

We love scraping, don&amp;#039;t we? But sometimes, we face Cloudflare protection. This script is designed to bypass the Cloudflare protection on websites, allowing you to interact with them programmatically.

- [Pour ceux qui voudrai remplacer FlareSolverr @ r/yggTorrents](https://www.reddit.com/r/yggTorrents/comments/1g7lsls/pour_ceux_qui_voudrai_remplacer_flaresolverr/).]]>
            </summary>
            <updated>2025-08-29T00:45:37+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3374</id>
            <title type="text"><![CDATA[Pipet]]></title>
            <link rel="alternate" href="https://github.com/bjesus/pipet" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3374"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[a swiss-army tool for scraping and extracting data from online assets, made for hackers.

Pipet is a command line based web scraper. It supports 3 modes of operation - HTML parsing, JSON parsing, and client-side JavaScript evaluation. It relies heavily on existing tools like curl, and it uses unix pipes for extending its built-in capabilities.]]>
            </summary>
            <updated>2025-08-29T01:18:57+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3666</id>
            <title type="text"><![CDATA[Firecrawl]]></title>
            <link rel="alternate" href="https://www.firecrawl.dev/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3666"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Turn websites into LLM-ready data.

Power your AI apps with clean data crawled from any website. It&amp;#039;s also open-source.
 🔥 Turn entire websites into LLM-ready markdown or structured data. Scrape, crawl and extract with a single API. 

- [Firecrawl @ GitHub](https://github.com/mendableai/firecrawl).

Related contents:

- [Firecrawl @ GitHub](https://github.com/firecrawl/firecrawl).
- [Firecrawl Observer @ GitHub](https://github.com/firecrawl/firecrawl-observer).

Related contents:

- [🚨 Someone built a tool that turns any website into clean data your AI can actually use @ Nav Toor&amp;#039;s X](https://nitter.net/heynavtoor/status/2031626457110425760).
- [Hermes Agent : veille technique auto-hébergée avec Matrix, FreshRSS et Firecrawl @ Cryptolab :fr:](https://cryptolab.re/posts/2026/hermes-agent-framework-self-hosted/).]]>
            </summary>
            <updated>2026-06-23T05:55:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/3883</id>
            <title type="text"><![CDATA[Scraperr]]></title>
            <link rel="alternate" href="https://github.com/jaypyles/Scraperr" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/3883"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Self-hosted webscraper.

Scraperr is a self-hosted web application that allows users to scrape data from web pages by specifying elements via XPath. Users can submit URLs and the corresponding elements to be scraped, and the results will be displayed in a table.]]>
            </summary>
            <updated>2025-08-29T02:43:41+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4135</id>
            <title type="text"><![CDATA[Browserless]]></title>
            <link rel="alternate" href="https://www.browserless.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4135"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A Pool of Hosted Browsers, For Use With Puppeteer or Playwright.

Run your scraping, testing, screenshotting or any other automation with our pool of browsers. Ready connect to with Puppeteer, Playwright or via our APIs.

- [Browserless @ GitHub](https://github.com/browserless/browserless).]]>
            </summary>
            <updated>2025-08-29T03:26:07+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4408</id>
            <title type="text"><![CDATA[monolith]]></title>
            <link rel="alternate" href="https://crates.io/crates/monolith" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4408"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[⬛️ CLI tool for saving complete web pages as a single HTML file.

A data hoarder’s dream come true: bundle any web page into a single HTML file. You can finally replace that gazillion of open tabs with a gazillion of .html files stored somewhere on your precious little drive.

- [monolith @ GitHub](https://github.com/Y2Z/monolith).
- [Monolith – L’outil parfait pour sauvegarder le web @ Korben :fr:](https://korben.info/monolith-archivage-web-html-autonome.html).]]>
            </summary>
            <updated>2025-08-29T04:11:26+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4612</id>
            <title type="text"><![CDATA[Dark Visitors]]></title>
            <link rel="alternate" href="https://darkvisitors.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4612"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A List of Known AI Agents on the Internet.

Insight into the hidden ecosystem of autonomous chatbots and data scrapers crawling across the web. Protect your website from unwanted AI agent access.

Related contents:

- [Comment protéger vos serveurs et lutter efficacement contre les crawlers d’IA @ Bearstech :fr:](https://bearstech.com/societe/blog/comment-proteger-vos-serveurs-et-lutter-efficacement-contre-les-crawlers-dia).]]>
            </summary>
            <updated>2025-10-30T06:47:32+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4671</id>
            <title type="text"><![CDATA[ScrapedIn]]></title>
            <link rel="alternate" href="https://github.com/dchrastil/ScrapedIn" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4671"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A tool to scrape LinkedIn without API restrictions for data reconnaissance.

This tool assists in performing reconnaissance using the LinkedIn.com website/API for red team or social engineering engagements. It performs a company specific search to extract a detailed list of employees who work for the target company. Enter the name of the target company and the tool will help determine the LinkedIn company ID, which will be used to perform the search.

- [🔗🤖 L&amp;#039;Intersection de l&amp;#039;ingénierie sociale et des outils de Scraping : découverte de ScrapedIn 🔍🌐 @ Serge Houtain&amp;#039;s LinkedIn :fr:](https://www.linkedin.com/posts/serge-houtain_github-dchrastilscrapedin-a-tool-to-scrape-activity-7139980953430450176-weIt/).]]>
            </summary>
            <updated>2025-08-29T04:54:57+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4741</id>
            <title type="text"><![CDATA[RSS-Bridge]]></title>
            <link rel="alternate" href="https://rss-bridge.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4741"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[RSS-Bridge is a PHP project capable of generating RSS and Atom feeds for websites that don&amp;#039;t have one. It can be used on webservers or as a stand-alone application in CLI mode. 

- [RSS-Bridge @ GitHub](https://github.com/RSS-Bridge/rss-bridge/).]]>
            </summary>
            <updated>2025-08-29T05:07:03+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5399</id>
            <title type="text"><![CDATA[Trafilatura]]></title>
            <link rel="alternate" href="https://trafilatura.readthedocs.io/en/latest/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5399"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A Python package &amp;amp; command-line tool to gather text on the Web.

Trafilatura is a Python package and command-line tool designed to gather text on the Web. It includes discovery, extraction and text processing components. Its main applications are web crawling, downloads, scraping, and extraction of main texts, metadata and comments. It aims at staying handy and modular: no database is required, the output can be converted to various commonly used formats.

- [Trafilatura @ GitHub](https://github.com/adbar/trafilatura).

Related contents:

- [Alimenter les RAG/LLM avec Trafilatura @ DevSecOps :fr:](https://blog.stephane-robert.info/docs/developper/programmation/python/trafilatura/).]]>
            </summary>
            <updated>2025-08-29T06:56:50+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5992</id>
            <title type="text"><![CDATA[htmlq]]></title>
            <link rel="alternate" href="https://github.com/mgdm/htmlq" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5992"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Like jq, but for HTML. Uses CSS selectors to extract bits of content from HTML files.]]>
            </summary>
            <updated>2025-08-29T08:36:45+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6168</id>
            <title type="text"><![CDATA[WebPlotDigitizer]]></title>
            <link rel="alternate" href="https://automeris.io/WebPlotDigitizer/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6168"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Extract data from plots, images, and maps.
A web based tool to extract numerical data from plot images. Supports XY, Polar, Ternary diagrams and Maps. 
It is often necessary to reverse engineer images of data visualizations to extract the underlying numerical data. WebPlotDigitizer is a semi-automated tool that makes this process extremely easy.

[WebPlotDigitizer @ GitHub](https://github.com/ankitrohatgi/WebPlotDigitizer)]]>
            </summary>
            <updated>2025-08-29T09:05:01+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6206</id>
            <title type="text"><![CDATA[Web Scraping Reference: Cheat Sheet for Web Scraping using R]]></title>
            <link rel="alternate" href="https://github.com/yusuzech/r-web-scraping-cheat-sheet" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6206"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Guide, reference and cheatsheet on web scraping using rvest, httr and Rselenium.
Inspired by Hartley Brody, this cheat sheet is about web scraping using rvest,httr and Rselenium. It covers many topics in this blog.

While Hartley uses python&amp;#039;s requests and beautifulsoup libraries, this cheat sheet covers the usage of httr and rvest. While rvest is good enough for many scraping tasks, httr is required for more advanced techniques. Usage of Rselenium(web driver) is also covered.]]>
            </summary>
            <updated>2025-08-29T09:11:02+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6376</id>
            <title type="text"><![CDATA[MrScraper]]></title>
            <link rel="alternate" href="https://mrscraper.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6376"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A hassle-free web scraper to process information from websites, easily and without getting blocked.]]>
            </summary>
            <updated>2025-08-29T09:41:17+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6681</id>
            <title type="text"><![CDATA[Browserflow]]></title>
            <link rel="alternate" href="https://browserflow.app/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6681"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Web Scraping &amp;amp; Web Automation.
Scrape websites. Automate tasks. No coding required.
Browserflow is a no-code/low-code Chrome extension that allows you to automate your work on any website. 
Save time by automating repetitive tasks in minutes. Run in your browser or in the cloud.]]>
            </summary>
            <updated>2025-08-29T10:30:44+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6945</id>
            <title type="text"><![CDATA[Buster]]></title>
            <link rel="alternate" href="https://github.com/dessant/buster" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6945"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Captcha solver extension for humans.
Buster is a browser extension which helps you to solve difficult captchas by completing reCAPTCHA audio challenges using speech recognition. Challenges are solved by clicking on the extension button at the bottom of the reCAPTCHA widget.

Related contents:

- [How I block all online ads @ Troubled Engineer](https://troubled.engineer/posts/no-ads/).]]>
            </summary>
            <updated>2025-12-09T07:25:28+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7091</id>
            <title type="text"><![CDATA[Txtpaper]]></title>
            <link rel="alternate" href="https://txtpaper.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7091"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Convert web pages into PDF, ePub, and Kindle (mobi) files]]>
            </summary>
            <updated>2025-08-29T11:39:19+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7226</id>
            <title type="text"><![CDATA[DaProfiler]]></title>
            <link rel="alternate" href="https://github.com/TheRealDalunacrobate/DaProfiler" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7226"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[DaProfiler allows you to get emails, social medias, adresses, works and more on your target using web scraping and google dorking techniques, based in France Only. The particularity of this program is its ability to find your targets e-mail adresses.]]>
            </summary>
            <updated>2025-08-29T12:02:32+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7415</id>
            <title type="text"><![CDATA[Linked Data Fragments]]></title>
            <link rel="alternate" href="http://linkeddatafragments.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7415"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Query the Web of data on Web-scale by
moving intelligence from servers to clients.]]>
            </summary>
            <updated>2025-08-29T12:34:49+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7624</id>
            <title type="text"><![CDATA[RoboBrowser]]></title>
            <link rel="alternate" href="https://github.com/jmcarp/robobrowser" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7624"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[RoboBrowser is a simple, Pythonic library for browsing the web without a standalone web browser. RoboBrowser can fetch a page, click on links and buttons, and fill out and submit forms. If you need to interact with web services that don&amp;#039;t have APIs, RoboBrowser can help.]]>
            </summary>
            <updated>2025-08-29T13:08:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/7722</id>
            <title type="text"><![CDATA[Portia | Scrapinghub]]></title>
            <link rel="alternate" href="http://scrapinghub.com/portia" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/7722"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Scrape websites visually. No code required!]]>
            </summary>
            <updated>2025-08-29T13:24:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8043</id>
            <title type="text"><![CDATA[Osmosis]]></title>
            <link rel="alternate" href="https://github.com/rc0x03/node-osmosis" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8043"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[HTML/XML parser and web scraper for NodeJS.]]>
            </summary>
            <updated>2025-08-29T14:17:51+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8206</id>
            <title type="text"><![CDATA[theharvester - Information Gathering]]></title>
            <link rel="alternate" href="https://code.google.com/p/theharvester" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8206"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[The objective of this program is to gather emails, subdomains, hosts, employee names, open ports and banners from different public sources like search engines, PGP key servers and SHODAN computer database.]]>
            </summary>
            <updated>2025-08-29T14:45:07+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8297</id>
            <title type="text"><![CDATA[Portia]]></title>
            <link rel="alternate" href="https://github.com/scrapinghub/portia" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8297"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Portia is a tool for visually scraping web sites without any programming knowledge. Just annotate web pages with a point and click editor to indicate what data you want to extract, and portia will learn how to scrape similar pages from the site.]]>
            </summary>
            <updated>2025-08-29T15:00:21+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/8764</id>
            <title type="text"><![CDATA[Scrapy]]></title>
            <link rel="alternate" href="http://scrapy.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/8764"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[An open source web scraping framework for Python.

Scrapy is a fast high-level screen scraping and web crawling framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

- [Scrapy @ GitHub](https://github.com/scrapy/scrapy)]]>
            </summary>
            <updated>2025-08-29T16:18:56+00:00</updated>
        </entry>
    </feed>
