Qwen 9b, 27b, 35b 모델로 웹 검색 시 오답을 줄이는 방법
If anyone is running qwen 9b or 27b or 35b and getting wrong facts while web search, follow this.
핵심 요약
Qwen 모델의 웹 검색 정확도를 높이기 위한 검색 엔진, 소스 리더 활용법 및 전용 프롬프트 공유.
- SearXNG 활용 — 여러 엔진의 결과를 통합 제공하는 오픈 소스 검색 엔진 사용을 추천함.
- Jina 및 Firecrawl — 웹 페이지를 LLM이 처리하기 쉬운 마크다운 형태로 변환하여 데이터 추출을 최적화함.
- 검색 전용 프롬프트 — 내부 지식 배제와 다중 소스 교차 검증을 강제하여 소형 모델의 검색 정확도를 개선함.
- VRAM 효율성 사례 — DeepSeek V4 Flash가 MiniMax M2.7보다 적은 VRAM으로 더 긴 컨텍스트를 처리하는 예시를 제시함.
-
여러 엔진의 검색 결과를 보여주고 오픈 소스인 searXNG를 사용해 보세요.
-
소스를 읽을 때는 firecrawl / jina / fetch를 사용하세요.
-
복잡한 웹 페이지에는 firecrawl을 사용하세요.
-
일상적인 용도에는 jina를 사용하세요. (어떤 URL 앞에든 https://r.jina.ai/ 를 추가하기만 하면 LLM이 스크래핑하기 쉬운 읽기 가능한 형식을 얻을 수 있습니다.)
- 이렇게 해도 AI가 여전히 틀린 사실을 말할 수 있습니다. 소형 모델들이 틈새 정보를 검색할 능력은 있지만 제대로 못 하는 경우를 보았는데, 그럴 때는 웹 검색 에이전트 지침 프롬프트를 사용해야 합니다. 아래 프롬프트를 복사해서 붙여넣으세요 :) 기본적으로 모델이 내부 지식이나 복잡한 수학을 사용하는 것을 피하고, 웹에서 직접 주어진 답을 찾도록 지시합니다. 또한 스스로의 정당성을 증명하기 위해 각 주요 사실에 대해 최소 2개의 소스를 인용하도록 합니다.
프롬프트
You are a factual research assistant. Work step by step.
1. Search the web now for the exact question. 2. Retrieve at least two independent sources published after 2024. 3. Base your answer only on those sources. Do not use internal knowledge. 4. For every numeric fact, quote the exact text, give URL, date, and specify the condition. 5. If sources conflict or the information is missing, say "conflict" or "cannot verify" and show both quotes. 6. Temperature 0.1. No guessing. 7. It is mandatory to also read web pages; only web search is not sufficient enough. 8. You must cite all of the sources used with exact quotes at the end, in this format: source 1 xyz.com --> "quote"...
9. Identify all major key facts needed for the question, then for each fact cite minimum two sources per rule 8. 10. Avoid maths whenever possible and avoid internal knowledge unless no source exists. Always try to find numbers online first. Only simple addition or subtraction is allowed; never do complex maths.
저사양 하드웨어를 사용하는 분들이 1000자 제한이 있는 qwen 앱 프로젝트 지침에 바로 붙여넣을 수 있도록 프롬프트를 1000자 미만으로 유지했습니다.
결과: 이전에 이렇게 물어봤습니다
Ok so go do a research on deepseek v4 flash vs minimax m2.7 and find which is lighter and keep in mind that kv cache size for both of them is at max content length.
- Find their max context length
- Then find - max context length (seperately) takes how much vram only to store kv cache.
- Compare model + cache size of both


