Executive TL;DR:
- Native MiniMax-H3 inference for Apple Silicon is being explored.
- Users report significant speed issues with current implementations.
- Sparse attention support could be a game-changer for speed.
The Buzz Score
The Internet’s Verdict: 70% Hyped, 30% Skeptical
H3-Metal: What’s the Fuss About?
Users are excited about the potential of H3-metal for Apple Silicon, but speed is a major concern. As one user noted:
I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest. The main issue is speed, a ~9-second 480×864 clip at 20 steps takes me a bit over an hour.
Looking for Speed Improvements
Some users are looking forward to improvements in speed. Another user said:
In the AMA Minimax said that H3 could support sparse attention, that would be a huge speedup! I wonder if there are any news on that. H3 is very cool. EDIT: testing a –sparse-attention optional mode based on what they said in the Reddit post.
Focus Keyword: H3 Metal