<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Vllm on wkqcosoft's Blog</title><link>https://wkqco33.github.io/wkqcosoft-docs/tags/vllm/</link><description>Recent content in Vllm on wkqcosoft's Blog</description><generator>Hugo</generator><language>ko-kr</language><copyright>Copyright @2024. wkqcosoft. All rights reserved.</copyright><lastBuildDate>Fri, 07 Aug 2026 22:37:01 +0900</lastBuildDate><atom:link href="https://wkqco33.github.io/wkqcosoft-docs/tags/vllm/index.xml" rel="self" type="application/rss+xml"/><item><title>vLLM을 활용한 로컬 SLM(소형 언어 모델) 고성능 서빙 완벽 가이드</title><link>https://wkqco33.github.io/wkqcosoft-docs/reports/2026-08-07-vllm%EC%9D%84-%ED%99%9C%EC%9A%A9%ED%95%B4-%EB%A1%9C%EC%BB%AC%EC%97%90%EC%84%9C-slm-%EC%86%8C%ED%98%95-%EC%96%B8%EC%96%B4-%EB%AA%A8%EB%8D%B8-%EC%84%9C%EB%B9%99-c3d24c/</link><pubDate>Fri, 07 Aug 2026 22:37:01 +0900</pubDate><guid>https://wkqco33.github.io/wkqcosoft-docs/reports/2026-08-07-vllm%EC%9D%84-%ED%99%9C%EC%9A%A9%ED%95%B4-%EB%A1%9C%EC%BB%AC%EC%97%90%EC%84%9C-slm-%EC%86%8C%ED%98%95-%EC%96%B8%EC%96%B4-%EB%AA%A8%EB%8D%B8-%EC%84%9C%EB%B9%99-c3d24c/</guid><description>vLLM 프레임워크를 이용해 로컬 환경에서 소형 언어 모델(SLM)을 고속 서빙하는 방법: 설치법, 모델 선정 기준, PagedAttention 및 연속 배치 최적화, OpenAI 호환 API 구축과 Ollama 비교</description></item></channel></rss>