Driving innovations with open communities. 💬 Join our Discord: discord.gg/hgKcfgXHAQ

HangZhou, China
Two LoRAs, one clean layer-editing workflow: extract an object as a transparent layer, then remove it from the original image. 🤖 modelscope.ai/models/DiffSyn… 🤖 modelscope.ai/models/DiffSyn… ✂️ LayerExtract isolates a prompt-specified subject and exports it on a transparent background for independent editing. ✨ LayerRemove takes the source image and a reference layer, removes the corresponding object, and reconstructs the scene behind it. ⚡ Both are compact Qwen-Image-2.1 LoRAs that can be hot-swapped within the same DiffSynth-Studio pipeline. 🧩 Together they turn flattened images into editable assets for compositing, repositioning, replacement, and design workflows. 📜 LoRA weights: Apache 2.0. Qwen-Image-2.1 base-model terms also apply.
1
2
18
988
Jina-OCR-v1 parses entire pages into structured Markdown at 2.57 pages per second.🚀 🤖 modelscope.ai/models/jinaai/… 📄 modelscope.ai/papers/2609.03… 🏆 Scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter. ⚡ Delivers the highest throughput among 14 evaluated systems at concurrency 32. 📑 Preserves text, formulas, tables, and reading order in a single pass. 🧠 The 3.4B MoE activates only 570M parameters per token, while FastMTP accelerates decoding without changing the autoregressive output. 📜 CC BY-NC 4.0. Commercial use requires permission.
1
4
24
1,834
NeoHorse-Jev-4B is now open—a compact model built to turn application states directly into structured decisions and probabilities. 🤖 modelscope.ai/collections/To… 🏆 Scores 77.70 across six text decision benchmark groups, ranking first among the four open-weight models with complete results in the comparison. ⚡ Prefill-only inference avoids autoregressive answer generation and supports three decision primitives: Choice for action selection, Noul for yes/no judgment, and Score for ordered ratings. 🎮 The same interface can power request routing, tool selection, workflow control, games, robot manipulation, and autonomous-driving simulations. 🖼️ Supports text inputs or a single image combined with text, returning normalized probabilities that applications can use directly. 🛠️ Deploy locally through vLLM, SGLang, Python, CLI, or HTTP. Apache 2.0.
21
41
351
19,146
Eight controls and inpainting in one checkpoint. Qwen-Image-2.1-Fun-Controlnet-Union adds precise structural guidance to Qwen-Image 2.1. 🎨 🤖 modelscope.ai/models/pai/Qwe… 🎛️ One control branch supports Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose, and Scribble with no checkpoint switching. 🖌️ Control and inpainting share the same branch, allowing masked regions to follow both the prompt and a structural reference. 🧠 Sixteen injection points guide every second Transformer block while keeping the base model frozen. ⚡ CFG-distilled sampling runs at guidance scale 1.0, while prefix KV caching accelerates repeated denoising steps. 📜 Qwen Research License. Qwen-Image 2.1 base weights are required.
3
4
71
5,172
NVIDIA Nemotron 3 Diarization is now available on ModelScope—adding live speaker attribution to existing ASR workflows without replacing the transcription model. 🎙️ 🤖 modelscope.ai/models/nv-comm… 👥 Processes streaming audio and returns speaker labels and timestamps for up to eight speaker slots in a single conversation. ⚡ Its end-to-end streaming architecture avoids separately combining voice activity detection, speaker embeddings, clustering, and post-processing. 🧠 The 99.2M-parameter model uses a 31-layer Transformer encoder with RoPE and builds on NVIDIA’s Streaming Sortformer architecture. 🔌 Pair it with Nemotron ASR, Parakeet, Canary, Whisper, or another ASR system to create speaker-attributed transcripts. 🏢 Designed for meetings, contact centers, clinical conversations, live captioning, media analysis, and multi-party voice agents. 🖥️ Supports NVIDIA Ampere, Hopper, and Blackwell GPUs, with inference through NeMo Speech C++.
Made with AI
3
10
71
4,895
Digital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. 📜 Apache 2.0 License. 🤖 modelscope.ai/models/XingChe… 📃 modelscope.ai/papers/2608.12… 🏆 Scores 96.87 overall on OmniDocBench v1.6, the highest among the listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. 📷 Handles digital, photographed, curved, and degraded documents directly, without a separate dewarping model. 🧠 Combines geometry-aware synthesis, consensus-generated labels, image-based self-verification, and progressive training from vision-language alignment to reinforcement learning. ⚡ Supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.
2
8
107
6,757
Xiaomi MiMo-V2.6 is now open—a native multimodal agent family built for large-scale reinforcement learning. 🚀📜 MIT License. 🤖 modelscope.cn/collections/Xi… 🏆 MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index and reaches 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. 🧠 The 1.02T MoE activates 42B parameters and supports text, images, video, audio, and a 1M-token context. ⚙️ One mixed RL run trains coding, general, visual, and cybersecurity agents together. Pro and Flash completed 30 steps each in under six days, producing around 750K trajectories. 🌐 The models support computer use, 3D creation, embodied control, coding, design, video, and music workflows.
6
18
208
9,118
Shanghai AI Lab and SJTU’s LUMIA Lab release NCP-ArchPreview, an open-weight 8.9B language modell. 📜 Apache 2.0. 🤖 modelscope.ai/collections/Sh… 📄 modelscope.ai/papers/2609.10… ⚡ Trained on 5.73T Dolma 3 tokens, it reaches OLMo-3-7B’s final Stage 1 loss with only 51.3% of the tokens, a 1.95× convergence gain. 🏆 Its Stage 1 macro-average rises from 46.59 to 49.04, with +5.99 on GSM8K and +4.28 on HumanEval. 🧠 NCP jointly predicts tokens and a concept sequence at one-quarter the length, then feeds those concepts back to guide generation. 🛠 Domain adaptation updates only the 17M-parameter concept module while keeping the token backbone frozen. 🚀 Concept-conditioned drafting improves mean accepted length by 4.17% with negligible overhead.
2
5
48
6,942
inclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows.📜 MIT License. 🤖 modelscope.ai/models/inclusi… 🤖 modelscope.ai/models/inclusi… 🏆 Design ranks #1 among open-weight models on the Artificial Analysis UI/UX Design leaderboard. Layer leads all 12 evaluated Crello settings and runs 4.3× faster than the evaluated 20B open-weight baseline. 🎨 Design generates complete UIs, dashboards, infographics, and posters at up to 2048×2048, with strong text rendering and native transparent RGBA output. 🧩 Layer decomposes flattened graphics into independently editable RGBA layers while preserving the original aspect ratio. 🧠 The pipeline combines multimodal prompt conditioning, a diffusion transformer, and a 4-channel VAE. Layer adds multi-frame generation for separate layer outputs.
3
7
83
4,038
Apsara 2026 is ON 🔥 Come find the ModelScope booth, snap a pic, and take home some merch: 🎒 Backpacks · Crossbody bags · Tote bags 🛋️ Neck pillows · ☕ Mugs · 💻 Laptop stands 🧸 Plush pendants · 🧢 Hats · 🧲 Fridge magnets 🎧 Earphone pouches · 📱 Phone chains ...and more. 📍 杭州国际博览中心二期 · 1F · 智能馆 · 1-5C · 魔搭社区 Hangzhou International Expo Center Phase II · 1F · Intelligence Engine · Booth 1-5C · ModelScope Come early before they're gone 👀
19
1,704
Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web, spanning domains such as productivity, development, graphics, multimedia, and science. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass both programmatic tests and VLM-based visual evaluation. The playground is ready. Let’s build! 🚀Dataset: modelscope.ai/datasets/Qwen/…
12
9
88
25,598
ModelScope now supports both inference and LoRA training for Qwen-Image-2.1 in the Civision.🎉 Jump in and start creating!👉modelscope.ai/civision
5
3
95
6,551
Introducing Qwen-Image-2.1: image generation and editing in one model, with native transparency and a compact 7B visual generation component.🚀 🤖 modelscope.ai/models/Qwen/Qw… 🎨 Try it in Studio: modelscope.ai/studios/Qwen/Q… ⚡ Efficient inference: KV cache reuse speeds up generation and editing while reducing memory usage, especially with multiple reference images. 🖼️ Native transparency: Generate regular or transparent images, edit transparent layers, and extract subjects from photos. 🛠️ Versatile editing: Combine up to 10 reference images, target local edits, and preserve portrait identity and product details. ✨ Refined aesthetics: Improved typography, portrait lighting, and realistic textures bring finer detail to generated images.
12
48
392
176,296
🎨 Qwen-Image-2.1 inference and training are now supported in ModelScope Civision, currently available on the CN site, with international support coming soon! 🤖 modelscope.cn/aigc
6
3,956
NetEase-Youdao releases Confucius4-R2T2, a true-streaming ASR model for live captions, simultaneous translation, and voice agents. 🤖 modelscope.ai/models/netease… 🏆 R2T2 achieves SOTA latency and recognition quality among the evaluated open-source models, while remaining competitive with leading closed-source systems. ⚡ Configurable 80ms–2s chunks deliver 200–600ms average latency with near-offline recognition accuracy. 📝 Append-only decoding commits stable text without revising earlier words, avoiding transcript flicker and giving downstream agents reliable input. 🌍 Optimized for Chinese and English, with multilingual recognition, hotword prompts, and contextual prompts. 🧰 The GitHub repo includes inference code, a minimal example, and a vLLM backend for both offline and real-time streaming ASR. 📜 Code: Apache 2.0. Weights: NetEase Model Use License Agreement.
2
9
47
4,483
Qwen-Image-2.1 is going open source! 🎨 ⌛Join the countdown modelscope.cn/models/Qwen/Qw… Weights + code. Download it, run it, build with it. 💻
29
114
932
206,561
🎨 Qwen-Image-2.1 is coming, and we're opening 50 early access spots for experienced creators and developers! 🔗 Apply here: alidocs.dingtalk.com/notable… 📮 We'll reach out by email if you're in. 💡 Program requirement: publish at least one original showcase or a hands-on review on your social media by Sep 28 at 23:59 (UTC+8). Your honest take, whether glowing or critical, is exactly what helps us make it better.
Made with AI
34
26
311
194,891
The early access emails are out! Check your inbox and enjoy! 🌟
2
1,425