Agnes-AI/Agnes-3.0-Flash 33B 멀티모달, AA 점수: 36
Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36
핵심 요약
262k 컨텍스트 윈도우와 하이브리드 어텐션 구조를 갖춘 33B 멀티모달 모델 Agnes-3.0-Flash가 공개되었습니다.
- 모델 구조 — 72개 레이어 중 54개는 재귀적 델타 규칙, 18개는 글로벌 어텐션을 사용하는 하이브리드 방식임.
- 멀티모달 성능 — 텍스트, 이미지, 비디오 이해가 가능하며 262,144 토큰의 긴 컨텍스트를 지원함.
- 오픈 모델 — Artificial Analysis에서 'Proprietary'로 분류되었으나 실제로는 오픈 모델로 공개됨.
- 추론 속도 — 'Flash'라는 이름에 걸맞게 빠른 속도를 목표로 설계됨.
huggingface.co
원문 사이트로 이동
HF에서 이 새로운 모델을 발견했다:
"까다로운 작업에 최적화됨. 262,144 토큰의 컨텍스트 윈도우, 조절 가능한 추론 노력, 도구 호출, 그리고 텍스트, 이미지, 비디오 이해 기능 탑재.
아키텍처
Agnes-3.0-Flash는 하이브리드 어텐션 디코더임: 4개 레이어 중 3개는 게이트 델타 규칙(시퀀스 길이와 무관하게 레이어별 상태를 유지하는 순환 방식)을 실행하고, 나머지 1개는 표준 글로벌 어텐션을 실행함. 따라서 72개 레이어 중 18개만 컨텍스트에 따라 커지는 KV 캐시를 가짐."
| Context length | 262 144 tokens |
|---|---|
| Decoder layers | 72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1 |
| Hidden size | 5120 |
| Global attention | 24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated output |
| Delta-rule layers | 16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32 |
| Feed-forward | SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layer |
| Positions | 3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims) |
| Vocabulary | 248 320 |
| Vision tower | 27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context. Context length 262 144 tokensDecoder layers 72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1Hidden size 5120Global attention 24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated outputDelta-rule layers 16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32Feed-forward SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layerPositions 3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)Vocabulary 248 320Vision tower 27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120 |

