[신규 모델] SupraLabs에서 Any2Any 모델 패밀리를 시작했습니다!
[NEW MODEL] SupraLabs started the Any2Any model family!
핵심 요약
텍스트, 이미지, 비디오를 단일 토큰으로 처리하는 30M 규모의 실험적 Any2Any 모델이 공개되었습니다.
- Any-to-Any 통합 — 텍스트, 이미지, 비디오를 별도의 인코더 없이 하나의 토큰 시퀀스로 처리함.
- 실험적 프로토타입 — 30M 파라미터 규모의 교육 및 개념 증명용 멀티모달 모델임.
- 단일 트랜스포머 구조 — 모달리티 간 경계 없이 다음 토큰을 예측하는 방식으로 멀티모달리티를 구현함.
- 기술적 한계 — 짧은 컨텍스트와 낮은 해상도 등 연구 목적의 초기 단계 모델임.
huggingface.co
원문 사이트로 이동
SupraLabs Supra-A2A-Nano-Exp - ~30M Any-to-Any Multimodal Transformer
Status: Experimental / Educational Prototype
Overview
Supra-A2A-Nano-Exp is a ~30M parameter autoregressive Transformer that unifies text, image, and video into a single token stream.
There are: - No separate vision encoder - No diffusion model - No cross-attention modules between modalities
Instead, everything is treated as tokens in one shared sequence.
Core Idea
The model predicts the next token in a unified stream where tokens can represent:
- Text tokens
- Image patches (VQ-VAE codes)
- Video frames (sequences of visual tokens)
👉 Multimodality = language modeling over a shared vocabulary.
Unified Token Stream Format
<TEXT>some text</TEXT> <IMAGE><FRAME>[64 visual tokens]</IMAGE> <VIDEO><FRAME>[frames of visual tokens]</VIDEO>
Tokenization
Text side
- GPT-2 BPE tokenizer: 50,257 tokens
- Special tokens (7):
<TEXT>,</TEXT><IMAGE>,</IMAGE><VIDEO>,</VIDEO><FRAME>
Total text vocab: 50,264 tokens
Vision side
- VQ-VAE encoder/decoder
- 3-layer convolutional encoder (/8 downsampling)
- Codebook: 256 entries × 64 dimensions
- Image 64×64 → 8×8 grid → 64 tokens
Combined vocabulary
50,264 (text) + 256 (visual) = 50,520 tokens
Architecture
| Component | Specification |
|---|---|


