이제 빠른 웹 리서치를 위해 클라우드 LLM을 쓸 필요가 없습니다
I no longer need a cloud LLM to do quick web research
핵심 요약
Qwen3.5와 MCP 도구를 활용해 로컬 환경에서 웹 검색 및 스크래핑 시스템을 구축한 사례 공유
- 로컬 모델 활용 — Qwen3.5:27B 모델을 RTX 4090에서 구동하여 클라우드 의존도를 낮춤
- MCP 서버 구축 — Playwright와 DuckDuckGo를 활용한 웹 스크래핑용 webmcp 서버를 공개함
- SearXNG 지원 — 커뮤니티 피드백을 반영하여 로컬 검색 엔진인 SearXNG 연동 기능을 즉시 추가함
- llama.cpp 연동 — llama.cpp Web UI의 MCP 도구 기능을 통해 로컬 LLM에 웹 브라우징 능력을 부여함
수정 2: SearXNG 지원이 추가되었습니다
어떤 분들에게는 아주 오래된 소식일 수도 있겠지만, 저는 최근에야 로컬 모델들이 제 품질 기준을 충족하기 시작해서 사용하기 시작했습니다. 제가 로컬에서 웹 검색/스크래핑을 위해 사용하고 있는 설정을 공유하고 싶습니다.
저는 RTX 4090에서 Qwen3.5:27B-Q3_K_M을 사용하며, 컨텍스트 길이는 약 200,000으로 설정했습니다. 속도는 약 40 tk/s 정도 나오고 VRAM은 약 22gb를 사용합니다.
llama.cpp Web UI를 통해 사용하며, MCP 도구를 활성화했습니다. 웹 검색/스크래핑을 위해 제공한 도구들은 다음과 같습니다:
"""
webmcp - 웹 스크래핑 및 콘텐츠 추출을 위한 MCP 서버
"""
import asyncio
import json
import logging
import os
import re
import time
from contextlib import contextmanager
from datetime import datetime, timezone
from pathlib import Path
from typing import Any
import httpx
from ddgs import DDGS
from markdownify import markdownify as md
from mcp.server.fastmcp import FastMCP
from mcp.server.transport_security import TransportSecuritySettings
from playwright.async_api import async_playwright
from readability import Document as ReadabilityDocument
from starlette.middleware.cors import CORSMiddleware
# ============================================================================
# Configuration
# ============================================================================
logger = logging.getLogger(__name__)
TOOL_CALL_LOG_PATH = os.path.join(
os.path.dirname(os.path.abspath(__file__)),
"tool_calls.log.json"
)
LLM_URL = os.environ.get("LLM_URL", "")
LLM_MODEL = os.environ.get("LLM_MODEL", "")
if not LLM_URL or not LLM_MODEL:
raise ValueError("LLM_URL and LLM_MODEL environment variables are required")
# ============================================================================
# Content Processing
# ============================================================================
def _html_to_clean(html: str) -> str:
"""HTML을 깨끗한 마크다운으로 변환하고 과도한 공백을 축소합니다."""
text = md(
html,
heading_style="ATX",
strip=["img", "script", "style", "nav", "footer", "header"]
)
# 3개 이상의 빈 줄을 2개로 축소
text = re.sub(r"\n{3,}", "\n\n", text)
# 각 줄의 공백(줄바꿈 제외) 축소
text = re.sub(r"[^\S\n]+", " ", text)
return text.strip()
async def _fetch_one(browser: Any, url: str, timeout_ms: int = 0) -> tuple[str, str]:
"""기존 브라우저 인스턴스를 사용하여 단일 URL을 가져옵니다."""
page = await browser.new_page()
await page.set_extra_http_headers({
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
})
try:
await page.goto(url, wait_until="domcontentloaded", timeout=timeout_ms)
await page.wait_for_timeout(2000)
html = await page.content()
finally:
await page.close()
doc = ReadabilityDocument(html)
title = doc.title()
clean_text = _html_to_clean(doc.summary())
if len(clean_text) < 50:
clean_text = _html_to_clean(html)
return title, clean_text
async def _fetch_pages(urls: list[str]) -> list[tuple[str, str, str | None]]:
"""공유 브라우저를 사용하여 여러 URL을 병렬로 가져옵니다. [(title, text, error)]를 반환합니다."""
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
try:
async def _fetch_single(url: str) -> tuple[str, str, str | None]:
try:
title, text = await _fetch_one(browser, url)
return title, text, None
except Exception as e:
logger.error(f"Failed to fetch {url}: {e}")
return "", "", str(e)
results = await asyncio.gather(*[_fetch_single(u) for u in urls])
finally:
await browser.close()
return results
async def _fetch_page_light(url: str) -> tuple[str, str]:
"""브라우저 없이 빠르게 가져오기 — 단순한 페이지에 적합합니다."""
async with httpx.AsyncClient(
timeout=30,
follow_redirects=True,
verify=False
) as client:
resp = await client.get(
url,
headers={"User-Agent": "Mozilla/5.0"}
)
resp.raise_for_status()
html = resp.text
doc = ReadabilityDocument(html)
title = doc.title()
clean_text = _html_to_clean(doc.summary())
if len(clean_text) < 50:
clean_text = _html_to_clean(html)
return


