AIchitect
StacksGraphBuilderSimulateCompareGenomeActivityPulse
AIchitect/Explore/Multimodal

Multimodal

11 tools
View in Explore graph →Open Builder →
OSSFreePopular

LLaVA

Open-source multimodal LLM assistant

⭐ 22,000
OSSFreePopular

Moondream

Tiny OSS vision language model

⭐ 11,000
OSSFree

Qwen-VL

Alibaba's open-weight vision-language model line (Qwen2.5-VL → Qwen3-VL)

⭐ 15,000
Free

Fal.ai

Fast serverless inference API for image, video, and audio models

⭐ 10,000
OSSFree

InternVL2

Top OSS multimodal model from OpenGVLab

⭐ 7,800
OSSFree

PaliGemma

Google's OSS vision-language model

⭐ 3,200

Pixtral

Mistral's vision-language model — folded into Mistral Small 4 (2026)

Free

Runway Gen-4.5

Frontier video generation with character/scene consistency

Free

Google Veo 3.1

Google's frontier video generation model with native audio

Free

Kling 3.0

Frontier video generation from Kuaishou

Seedance 2.5

ByteDance's frontier text/image-to-video model

Other categories

Coding Assistants (27)Autonomous Agents (11)Agent Frameworks (16)Pipelines & RAG (12)LLM Infrastructure (45)Design & UI (11)AI-Augmented Code Quality (6)Documentation (9)Product & PM (6)MCP Servers (20)Prompt & Eval (10)Specifications (11)Spec-Driven Dev (5)AI Code Review (4)Fine-tuning (7)Voice AI (11)Browser Automation (7)Observability (14)Memory & Persistence (7)AI Guardrails (6)