구글, Gemma 4 채팅 템플릿 업데이트: 툴 콜링 개선, '게으름' 감소, Hopper GPU용 Flash Attention 4 지원 및 비전 가이드 공개
Google is updating Gemma 4's chat templates, bringing major fixes to tool calling and reducing "laziness", and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!
핵심 요약
Gemma 4의 툴 콜링 오류와 모델의 '게으름' 문제를 해결하는 채팅 템플릿 업데이트가 배포되었습니다.
- 채팅 템플릿 업데이트 — 툴 콜링 오류 수정 및 모델의 응답 지연 현상 완화함
- 성능 최적화 — Hopper GPU에서 Flash Attention 4 지원 추가됨
- 비전 가이드 — Gemma 4 비전 모델 활용 및 개선을 위한 대화형 가이드 제공됨
- 기술적 해결책 — 템플릿 수정을 통해 모델의 추론 능력과 응답 품질 향상됨
...그리고 추론(thinking) 보존!!!!!!!!
이미지 링크는 무시하고 소스는 여기 있음: https://x.com/googlegemma/status/2077449152062247219
https://huggingface.co/spaces/google/gemma4_vision_token_budget

