JEV처럼 어떤 LLM이든 사용할 수 있음
You can use any LLM just like JEV
핵심 요약
LLM의 추론을 끄고 로짓 확률을 활용해 분류 작업을 빠르고 효율적으로 수행하는 방법 공유.
- GGUF 활용 — llama.cpp를 통해 LLM의 추론을 끄고 확률값만 추출함
- 분류 최적화 — 스팸 분류 등 이진 분류 작업을 매우 빠른 속도로 처리함
- 다중 질문 처리 — 비트 연산을 통해 한 번의 추론으로 여러 질문에 대한 확률을 계산함
- 성능 이점 — 디코딩 과정을 생략하여 일반적인 LLM보다 훨씬 빠른 응답 속도를 보임
You can simply run any GGUF with llama.cpp with n_predict=1 and n_probs=10, disable reasoning, and prompt it such as "If the following email is spam, respond with 1, if not spam, respond with 0. Do not respond with anything other than 1 or 0. Email: ...."
And that is it! It returns confidence percentages such as:
1 = 94.9%
0 = 5.08%
Example:
llama-server -m "C:\Users\MyUserName\llama.cpp\models\Spark-X2.5-4B-Q4_K_M.gguf" -c 4096 -ngl all -fit off -fa on -b 2048 -ub 512 -np 1 --cache-ram 0 --reasoning off --no-reasoning-preserve --perf
Then:
curl.exe -s -X POST http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d "{\"messages\":[{\"role\":\"system\",\"content\":\"Classify spam. Reply only 1=spam or 0=not spam.\"},{\"role\":\"user\",\"content\":\"CONGRATULATIONS!!! You have won $5,000,000! Click here immediately to claim your prize!\"}],\"max_tokens\":1,\"logprobs\":true,\"top_logprobs\":10,\"temperature\":1.0,\"top_p\":1.0}"
Result:
{"choices":[{"finish_reason":"length","index":0,"message":{"role":"assistant","content":"1"},"logprobs":{"content":[{"id":30,"token":"1","bytes":[49],"logprob":-0.00456317700445652,"top_logprobs":[{"id":30,"token":"1","bytes":[49],"logprob":-0.00456317700445652},{"id":29,"token":"0","bytes":[48],"logprob":-5.395024299621582},{"id":1033,"token":"**","bytes":[42,42],"logprob":-12.013711929321289},{"id":1046,"token":"The","bytes":[84,104,101],"logprob":-13.005236625671387},{"id":198,"token":"\n","bytes":[10],"logprob":-13.100714683532715},{"id":54,"token":"I","bytes":[73],"logprob":-14.624603271484375},{"id":3640,"token":"This","bytes":[84,104,105,115],"logprob":-14.800630569458008},{"id":130977,"token":"<tool_call>","bytes":[60,116,111,111,108,95,99,97,108,108,62],"logprob":-14.971238136291504},{"id":6908,"token":"Class","bytes":[67,108,97,115,115],"logprob":-15.373867988586426},{"id":3923,"token":"class","bytes":[99,108,97,115,115],"logprob":-15.442902565002441}]}]}}],"created":1789950066,"model":"C:\\Users\\MyUserName\\llama.cpp\\models\\Spark-X2.5-4B-Q4_K_M.gguf","system_fingerprint":"b11026-b49650adb","object":"chat.completion","usage":{"completion_tokens":1,"prompt_tokens":64,"total_tokens":65,"prompt_tokens_details":{"cached_tokens":59}},"id":"chatcmpl-x2WrCObzFNYjKVkwDmcL8FLquwfZ0NEa","timings":{"cache_n":59,"prompt_n":5,"prompt_ms":634.566,"prompt_per_token_ms":126.9132,"prompt_per_second":7.879401039450585,"predicted_n":1,"predicted_ms":0.001,"predicted_per_token_ms":0.0,"predicted_per_second":0.0}}

