<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>vllm</title>
    <link rel="self" type="application/atom+xml" href="https://links.biapy.com/guest/tags/3552/feed"/>
    <updated>2026-09-25T14:35:41+00:00</updated>
    <id>https://links.biapy.com/guest/tags/3552/feed</id>
            <entry>
            <id>https://links.biapy.com/links/14007</id>
            <title type="text"><![CDATA[Model Optimizer]]></title>
            <link rel="alternate" href="https://nvidia.github.io/Model-Optimizer/" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/14007"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

- [Model Optimizer @ GitHub](https://github.com/NVIDIA/Model-Optimizer).]]>
            </summary>
            <updated>2026-09-25T12:28:55+00:00</updated>
        </entry>
            <entry>
            <id>https://links.biapy.com/links/14005</id>
            <title type="text"><![CDATA[kvcached]]></title>
            <link rel="alternate" href="https://github.com/ovg-project/kvcached" />
            <link rel="via" type="application/atom+xml" href="https://links.biapy.com/links/14005"/>
            <author>
                <name><![CDATA[Biapy]]></name>
            </author>
            <summary type="text">
                <![CDATA[Make GPU Sharing Flexible and Easy .
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond.

kvcached (KV cache daemon) is a KV cache library for LLM serving/training on shared GPUs. By bringing OS-style virtual memory abstraction to LLM systems, it enables elastic and demand-driven KV cache allocation, improving GPU utilization under dynamic workloads.]]>
            </summary>
            <updated>2026-09-25T06:34:54+00:00</updated>
        </entry>
    </feed>
