학습 완료아이폰에서 20B MoE 모델 120 tok/s 구동 — Maple-Preview 데모2026-08-05 23:38 · HN (show hn) · 원문 보기 ↗ #on-device-llm#ai-inference#moe근거: AI/LLM 추론 최적화 관심 분야 — 아이폰에서 20B MoE 모델을 120 tok/s로 구동하는 on-device 추론 기술 사례액션: deepgrove.ai/maple-preview 페이지에서 ternary quantization 방식과 MoE 구조 확인, 관련 오픈소스(BitNet, llama.cpp iOS 브랜치 등) 비교 학습