🏗️ Building on HF
Adarsh Zolekar
adarshzolekar
AI & ML interests
Exploring AI, ML, Deep Learning, models and datasets while building and contributing to the Hugging Face community.
Recent Activity
liked a model 4 days ago
Lightricks/LTX-2.5 liked a model 4 days ago
Qwen/Qwen-Image-2.1 upvoted a paper 4 days ago
Grounded Skill Synthesis from Code at Scale for Agentic IntelligenceOrganizations
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 22.1M • 1.57k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 341k • • 978 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 609k • • 5.37k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 692k • • 15k
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 246M • • 6.12k -
BAAI/bge-m3
Sentence Similarity • Updated • 36.6M • • 3.65k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 13.9M • 940 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 17.1M • • 1.22k
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.6M • • 6.4k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.48M • • 3.4k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.7M • • 7.01k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.44M • 2k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.26M • 3.56k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 12.4M • • 2.05k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 3.82M • • 4.01k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 11.2M • 1.07k
Reasoning & Agentic Models
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 246M • • 6.12k -
BAAI/bge-m3
Sentence Similarity • Updated • 36.6M • • 3.65k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 13.9M • 940 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 17.1M • • 1.22k
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.6M • • 6.4k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.48M • • 3.4k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.7M • • 7.01k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.44M • 2k
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 22.1M • 1.57k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 341k • • 978 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 609k • • 5.37k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 692k • • 15k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.26M • 3.56k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 12.4M • • 2.05k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 3.82M • • 4.01k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 11.2M • 1.07k