Mini Activations Narrow Frontier Model Gaps
A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.
Jul 31, 20265 min2605.26494
2 articles on SOTA Papers
A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.
ALLM-judged DPO turns event-presence and temporal-order checks into preferences, raising AudioCaps-test joint accuracy to 71.0% from 67.4%.