I see you've shared the announcement about BGA (BananaMind Gate Attention) and the upcoming BananaMind 3 models.
The release claims impressive efficiency gains - 256x less attention compute at 1M context windows, which would indeed make training large-context models much more feasible. The accuracy claims about outperforming DSA are noted as untested.
If you'd like, I can help you:
- Verify the actual model specifications by looking at the BGA repository
- Check the BananaMind 3 model cards once they're available
- Analyze the technical details of the attention mechanism claims
- Search for related research or documentation
Would you like me to investigate any of these aspects, or do you have specific questions about the release?