텐센트의 Youtu-Parsing-Omni 모델 공개
tencent/Youtu-Parsing-Omni · Hugging Face
핵심 요약
문서, 이미지, 오디오, 비디오 등 다양한 입력을 하나의 JSON 구조로 파싱하는 5B 규모의 옴니 모달 모델입니다.
- 통합 파싱 스키마 — 작업 프롬프트를 통해 7가지 데이터 유형을 하나의 JSON으로 처리함
- 옴니 인코더 — 이미지, 오디오, 비디오 입력을 단일 모델에서 통합 처리함
- 고성능 경량 모델 — 5B 파라미터로 OmniDocBench 등에서 최상위권 성능을 기록함
- 간편한 서빙 — vLLM 플러그인과 예제 코드를 제공하여 배포가 용이함
huggingface.co
원문 사이트로 이동
Youtu-Parsing-Omni는 5B 파라미터의 콤팩트한 옴니 모달 파싱 모델임. 문서 페이지, 일반 이미지, 차트/순서도, 기하학 도형, 오디오 클립, 혹은 시청각 영상 등 어떤 입력이 들어오든 인식(레이아웃 요소, 텍스트, 표, 수식, 바운딩 박스, 타임스탬프, ASR, OCR, 음향 이벤트, 카메라 움직임)과 인지(캡션, 서술, 보고서)를 모두 아우르는 구조화된 JSON 결과물을 뱉어냄. 출력 형식은 작업 프롬프트(--task, prompts/youtu_parsing_omni.json의 키값)를 통해 선택 가능함.
| Input | --task | modality / subtype | Key contents |
|---|---|---|---|
| Document page | document | image / document | layout elements with bbox, text / LaTeX / OTSL tables / Markdown charts / Mermaid flowcharts, reading order |
| Natural image | natural_image | image / natural_image | entities and text with bbox, tags, captions, global description |
| Chart | graphics_chart | image / document | one chart element: Markdown table, notes, caption |
| Flowchart | graphics_flowchart | image / document | one flowchart element: Mermaid, caption |
| Geometry figure | graphics_geometric | image / document | one geometric element: points, lines, arcs, shapes, geometric relations and measurements |


