<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>nlp</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/134/feed"/>
    <updated>2026-08-02T01:35:37+00:00</updated>
    <id>https://links.biapy.com/guest/tags/134/feed</id>
            <entry>
            <id>https://links.biapy.com/links/11240</id>
            <title type="text"><![CDATA[Mark V. Shaney Junior Gibberish Generator]]></title>
            <link rel="alternate" href="https://github.com/susam/mvs" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/11240"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A minimum viable Markov gibberish generator in 32 lines of Python, inspired by the legendary Mark V. Shaney program of 1980s.

Related contents:

- [Fed 24 Years of My Blog Posts to a Markov Model @ Susam Pal](https://susam.net/fed-24-years-of-posts-to-markov-model.html).]]>
            </summary>
            <updated>2025-12-15T12:56:39+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/10501</id>
            <title type="text"><![CDATA[unified]]></title>
            <link rel="alternate" href="https://unifiedjs.com/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/10501"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[unified is a collective of 500+ free and open source packages that work with content as structured data (ASTs).
Different kinds of content can be connected together. Particularly, markdown, natural language, HTML, XML, and JavaScript are frequently used.

Content as structured data.

We compile content to syntax trees and syntax trees to content.
We also provide hundreds of packages to work on the trees in between.
You can build on the unified collective to make all kinds of interesting things.

- [unified @ GitHub](https://github.com/unifiedjs).

Related contents:

- [How to keep package.json under control @ val town](https://blog.val.town/gardening-dependencies).]]>
            </summary>
            <updated>2025-10-02T06:41:10+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/44</id>
            <title type="text"><![CDATA[spaCy]]></title>
            <link rel="alternate" href="https://spacy.io/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/44"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[💫 Industrial-strength Natural Language Processing (NLP) in Python.

spaCy is a library for advanced Natural Language Processing in Python and Cython. It&amp;#039;s built on the very latest research, and was designed from day one to be used in real products.

spaCy comes with pretrained pipelines and currently supports tokenization and training for 70+ languages. It features state-of-the-art speed and neural network models for tagging, parsing, named entity recognition, text classification and more, multi-task learning with pretrained transformers like BERT, as well as a production-ready training system and easy model packaging, deployment and workflow management. spaCy is commercial open-source software, released under the MIT license.

- [spaCy @ GitHub](https://github.com/explosion/spaCy).

Related contents:

- [Embedding Millions of Text Documents With Qwen3 @ daft](https://www.daft.ai/blog/embedding-millions-of-text-documents-with-qwen3).
- [Learn How to Use Transformers with HuggingFace and SpaCy @ towards data science](https://towardsdatascience.com/mastering-nlp-with-spacy-part-4/).]]>
            </summary>
            <updated>2025-09-18T05:52:31+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1460</id>
            <title type="text"><![CDATA[AntiSquat]]></title>
            <link rel="alternate" href="https://github.com/redhuntlabs/antisquat" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1460"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[AntiSquat leverages AI techniques such as natural language processing (NLP), large language models (ChatGPT) and more to empower detection of typosquatting and phishing domains.

Related contents:

- [ 🚨🚨 AntiSquat : l’IA qui traque les faux sites avant qu’ils ne vous piégent ! 🚨🚨 @ Laurent Biagotti&amp;#039;s LinkedIn :fr:](https://www.linkedin.com/posts/laurent-biagiotti-19779284_cybersaezcuritaez-phishing-typosquatting-activity-7298965818971684864-Whqm/).]]>
            </summary>
            <updated>2025-08-28T19:59:23+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1512</id>
            <title type="text"><![CDATA[Open Voice OS]]></title>
            <link rel="alternate" href="https://www.openvoiceos.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1512"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[🌟 OpenVoiceOS is an open-source platform for smart speakers and other voice-centric devices.

OpenVoiceOS is a community-driven, open-source voice AI platform for creating custom voice-controlled ​interfaces across devices with NLP, a customizable UI, and a focus on privacy and security.

- [Open Voice OS core @ GitHub](https://github.com/OpenVoiceOS/ovos-core).]]>
            </summary>
            <updated>2025-08-28T20:08:22+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/1849</id>
            <title type="text"><![CDATA[DataBridge]]></title>
            <link rel="alternate" href="https://databridge.gitbook.io/databridge-docs" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/1849"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Multi-modal modular data ingestion and retrieval.

DataBridge is an open source library for natural language search and management of multi-modal data. Get started by installing databridge now!

DataBridge is a powerful document processing and retrieval system designed for building intelligent document-based applications. It provides a robust foundation for semantic search, document processing, and AI-powered document interactions.]]>
            </summary>
            <updated>2025-08-28T21:04:04+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4214</id>
            <title type="text"><![CDATA[Dataherald AI]]></title>
            <link rel="alternate" href="https://dataherald.readthedocs.io/en/latest/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4214"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Dataherald is a natural language-to-SQL engine built for enterprise-level question answering over relational data. It allows you to set up an API from your database that can answer questions in plain English.]]>
            </summary>
            <updated>2025-08-29T03:39:09+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4413</id>
            <title type="text"><![CDATA[life2vec - Official Model and Paper Page]]></title>
            <link rel="alternate" href="https://life2vec.dk/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4413"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Using Sequences of Life-events to Predict Human Lives.

We represent human lives in a way that shares structural similarity to language, and we exploit this similarity to adapt natural language processing techniques to examine the evolution and predictability of human lives based on detailed event sequences. We do this by drawing on a comprehensive registry dataset, which is available for Denmark across several years, and that includes information about life-events related to health, education, occupation, income, address and working hours, recorded with day-to-day resolution. 

- [live2vec @ GitHub](https://github.com/SocialComplexityLab/life2vec).
- [Life2vec – Une IA danoise qui prédit votre vie et… votre mort ! @ Korben :fr:](https://korben.info/life2vec-ia-danoise-predit-vie-donnees.html).]]>
            </summary>
            <updated>2025-08-29T04:12:25+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4877</id>
            <title type="text"><![CDATA[The Grand Complete Data Science Guide With Videos And Materials]]></title>
            <link rel="alternate" href="https://github.com/krishnaik06/The-Grand-Complete-Data-Science-Materials" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4877"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Contribute to krishnaik06/The-Grand-Complete-Data-Science-Materials development by creating an account on GitHub.]]>
            </summary>
            <updated>2025-08-29T05:31:07+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4963</id>
            <title type="text"><![CDATA[Apache UIMA]]></title>
            <link rel="alternate" href="https://uima.apache.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4963"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Unstructured Information Management applications are software systems that analyze large volumes of unstructured information in order to discover knowledge that is relevant to an end user. An example UIM application might ingest plain text and identify entities, such as persons, places, organizations; or relations, such as works-for or located-at. 

UIMA enables applications to be decomposed into components, for example &amp;quot;language identification&amp;quot; =&amp;gt; &amp;quot;language specific segmentation&amp;quot; =&amp;gt; &amp;quot;sentence boundary detection&amp;quot; =&amp;gt; &amp;quot;entity detection (person/place names etc.)&amp;quot;.]]>
            </summary>
            <updated>2025-08-29T05:44:13+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4964</id>
            <title type="text"><![CDATA[GATE]]></title>
            <link rel="alternate" href="https://gate.ac.uk/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4964"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[General Architecture for Text Engineering.

GATE is an open source software toolkit capable of solving almost any text processing problem.]]>
            </summary>
            <updated>2025-08-29T05:44:13+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/4975</id>
            <title type="text"><![CDATA[txtai]]></title>
            <link rel="alternate" href="https://neuml.github.io/txtai/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/4975"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[txtai is an all-in-one embeddings database for semantic search, LLM orchestration and language model workflows.

[txtai @ GitHub](https://github.com/neuml/txtai).]]>
            </summary>
            <updated>2025-08-29T05:47:15+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5102</id>
            <title type="text"><![CDATA[Moses]]></title>
            <link rel="alternate" href="http://www2.statmt.org/moses/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5102"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Moses, the machine translation system.

Moses is a statistical machine translation system that allows you to automatically train translation models for any language pair. All you need is a collection of translated texts (parallel corpus). Once you have a trained model, an efficient search algorithm quickly finds the highest probability translation among the exponential number of choices. 

[Moses @ GitHub](https://github.com/moses-smt/mosesdecoder).]]>
            </summary>
            <updated>2025-08-29T06:07:29+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5103</id>
            <title type="text"><![CDATA[text2vec]]></title>
            <link rel="alternate" href="https://text2vec.org/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5103"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[text2vec is an R package which provides an efficient framework with a concise API for text analysis and natural language processing (NLP).

[text2vec @ GitHub](https://github.com/dselivanov/text2vec).]]>
            </summary>
            <updated>2025-08-29T06:07:29+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5104</id>
            <title type="text"><![CDATA[MITIE]]></title>
            <link rel="alternate" href="https://github.com/mit-nlp/MITIE" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5104"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[library and tools for information extraction.

This project provides free (even for commercial use) state-of-the-art information extraction tools. The current release includes tools for performing named entity extraction and binary relation detection as well as tools for training custom extractors and relation detectors.]]>
            </summary>
            <updated>2025-08-29T06:07:30+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/5502</id>
            <title type="text"><![CDATA[StableLM:]]></title>
            <link rel="alternate" href="https://github.com/stability-AI/stableLM/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/5502"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Stability AI Language Models.

This repository contains Stability AI&amp;#039;s ongoing development of the StableLM series of language models and will be continuously updated with new checkpoints. The following provides an overview of all currently available models. More coming soon.]]>
            </summary>
            <updated>2025-08-29T07:14:00+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6173</id>
            <title type="text"><![CDATA[NLP Cheat Sheet]]></title>
            <link rel="alternate" href="https://github.com/janlukasschroeder/nlp-cheat-sheet-python" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6173"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Introduction to Natural Language Processing (NLP) tools, frameworks, concepts, resources for Python]]>
            </summary>
            <updated>2025-08-29T09:05:59+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6182</id>
            <title type="text"><![CDATA[Advanced NLP with spaCy · A free online course]]></title>
            <link rel="alternate" href="https://course.spacy.io/en" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6182"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[spaCy is a modern Python library for industrial-strength Natural Language Processing. In this free and interactive online course, you&amp;#039;ll learn how to use spaCy to build advanced natural language understanding systems, using both rule-based and machine learning approaches.]]>
            </summary>
            <updated>2025-08-29T09:07:01+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/6630</id>
            <title type="text"><![CDATA[PaddleNLP]]></title>
            <link rel="alternate" href="https://github.com/PaddlePaddle/PaddleNLP" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/6630"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[PaddleNLP is an easy-to-use and powerful natural language processing development library. Aggregates high-quality pre-trained models in the industry and provides an out -of-the-box development experience. The model library covering multiple scenarios of NLP and industrial practice examples can meet the needs of developers for flexible customization .]]>
            </summary>
            <updated>2025-08-29T10:22:39+00:00</updated>
        </entry>
    </feed>
