<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>gemm</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/3456/feed"/>
    <updated>2026-07-21T16:11:52+00:00</updated>
    <id>https://links.biapy.com/guest/tags/3456/feed</id>
            <entry>
            <id>https://links.biapy.com/links/12584</id>
            <title type="text"><![CDATA[DeepGEMM]]></title>
            <link rel="alternate" href="https://github.com/deepseek-ai/DeepGEMM" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/12584"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling.

DeepGEMM is a unified, high-performance tensor core kernel library that brings together the key computation primitives of modern large language models — GEMMs (FP8, FP4, BF16), fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer, HyperConnection (HC), and more — into a single, cohesive CUDA codebase. All kernels are compiled at runtime via a lightweight Just-In-Time (JIT) module, requiring no CUDA compilation during installation.]]>
            </summary>
            <updated>2026-04-20T11:54:19+00:00</updated>
        </entry>
    </feed>
