Alibaba ने Qwen को पूर्ण agent workflows की ओर मोड़ा

Alibaba की Qwen टीम ने Qwen3.7-Plus जारी किया है, एक नया multimodal मॉडल जो visual understanding को coding और tool use जैसी पारंपरिक agent क्षमताओं के साथ जोड़ता है। कंपनी इसे एक multimodal interactive hybrid agent के रूप में वर्णित करती है, और इसकी positioning उल्लेखनीय है: इसे image input वाले chatbot के रूप में नहीं, बल्कि interfaces को समझकर उनके भीतर action लेने वाले system के रूप में पेश किया गया है।

प्रदान किए गए source text के अनुसार, Qwen3.7-Plus को वास्तविक दुनिया के दृश्यों को पहचानने, स्क्रीन सामग्री पढ़ने, graphical user interfaces संचालित करने, visual templates से code लिखने, और mobile apps को end to end navigate करने के लिए डिज़ाइन किया गया है। Operating model महत्वपूर्ण है। UI clicks और command-line instructions एक ही agent loop के भीतर चलते हैं, जो संकेत देता है कि Alibaba perception, planning, और execution के लिए अलग-अलग models के बजाय automation के अधिक unified रूप की ओर बढ़ रहा है।

लंबे समय तक चलने वाले tasks इस pitch के केंद्र में हैं

Alibaba के showcase examples extended workflows पर autonomy पर केंद्रित हैं। एक demonstration में, एक hybrid agent system ने 11 घंटे से अधिक समय में एक English vocabulary learning app बनाया। Source के अनुसार, इस run में 1,000 से अधिक agent calls के दौरान 10,000 से अधिक lines of code तैयार हुईं।

बताई गई प्रक्रिया में requirements documentation, automated code generation, dependency installation, test-case creation, GUI-based testing, parallel test scenarios, और version management शामिल थे। ये विवरण महत्वपूर्ण हैं क्योंकि वे कहानी को एक one-shot coding demo से आगे ले जाते हैं। Alibaba का तर्क है कि model एक multi-stage software project के दौरान लगातार काम कर सकता है और बार-बार human hand-holding के बिना tools और interfaces के बीच काम करता रह सकता है।

दूसरी demonstration software generation से software imitation की ओर गई। Alibaba का कहना है कि agent ने interface को parse करके, SwiftUI code बनाकर, बाहरी real-time stock data API जोड़कर, result compile करके, और अपने दम पर दस functional tests चलाकर Apple के native macOS Stocks app की नकल की। यदि यह प्रदर्शन व्यापक रूप से लागू होता है, तो model का मूल्य prompts का जवाब देने में कम और एक working interface को देखने और उसे code में reproduce करने के बीच का समय घटाने में अधिक हो सकता है।

Browser और cloud operations दायरा बढ़ाते हैं

तीसरा use case model को browser-based operations तक विस्तारित करता है। Qwen for Chrome नामक एक sidebar extension के माध्यम से, system user permission के साथ agent mode में बदल सकता है और cloud console tasks कर सकता है। Source text में एक उदाहरण का उल्लेख है जिसमें model ने उपलब्ध सबसे सस्ता virtual server instance खरीदा, जिसमें image, storage, और security-group options की setup भी शामिल थी।

Alibaba यह भी कहता है कि model ने बाद के scaling और maintenance tasks संभाले। यह महत्वपूर्ण है क्योंकि यह pitch को isolated task completion से lifecycle management की ओर ले जाता है। एक model जो service बना, test कर, configure कर, और बाद में maintain कर सकता है, वह उस क्षेत्र में प्रवेश करता है जिसे enterprises परंपरागत रूप से engineers, scripts, और workflow tools के संयोजन के लिए सुरक्षित रखते हैं।

मजबूत GUI performance, लेकिन pure reasoning कमजोर

प्रदान की गई सामग्री में benchmark picture मिश्रित है। Alibaba के प्रकाशित results के अनुसार Qwen3.7-Plus graphical interface tasks पर विशेष रूप से अच्छा प्रदर्शन करता है। AndroidWorld और ScreenSpot Pro पर, model को GPT-5.4 (xhigh) से काफी आगे बताया गया है। इससे Alibaba को एक crowded market में एक ठोस angle मिलता है: यदि interface manipulation AI का बड़ा battleground बनता है, तो Qwen केवल बातचीत नहीं, execution पर प्रतिस्पर्धा करना चाहता है।

साथ ही, source text कहता है कि system pure logic benchmarks में पिछड़ जाता है। यह caveat महत्वपूर्ण है। यह संकेत देता है कि Qwen3.7-Plus उन स्थितियों में अधिक उपयोगी हो सकता है जहाँ environment स्वयं structure, visual anchors, और action affordances देता है, बजाय उन abstract reasoning tasks के जहाँ model को उस context के बिना काम करना पड़े।

व्यावहारिक रूप से, model की ताकत software को देखकर और उसके भीतर action लेने में निहित दिखती है। यह intelligence की एक अधिक संकीर्ण, लेकिन व्यावसायिक रूप से प्रासंगिक परिभाषा है, खासकर enterprise automation, testing, customer operations, और software prototyping के लिए।

यह release क्यों मायने रखती है

Qwen3.7-Plus को Alibaba Cloud के माध्यम से एक proprietary लेकिन अपेक्षाकृत सस्ते option के रूप में भी position किया गया है। Price और deployment path महत्वपूर्ण हैं क्योंकि agentic systems लंबे sessions चलाने, कई calls execute करने, और external tools से interact करने पर जल्दी महंगे हो सकते हैं। यदि Alibaba strong interface performance देते हुए operating costs कम रख सकता है, तो उसे developers और businesses के बीच एक स्वीकार्य market मिल सकता है जो frontier-model pricing के बिना automation चाहते हैं।

इस release का व्यापक महत्व यह है कि Qwen3.7-Plus AI vendors की progress को परिभाषित करने के तरीके में बदलाव को दर्शाता है। केवल benchmark scores या chat quality पर ध्यान देने के बजाय, Alibaba इस बात पर जोर दे रहा है कि क्या model interface का निरीक्षण कर सकता है, निर्णय ले सकता है, tools call कर सकता है, code लिख सकता है, और घंटों तक task पर बना रह सकता है। इससे reliability, oversight, और failure handling से जुड़े कठिन सवाल हल नहीं हो जाते। लेकिन यह दिखाता है कि प्रतिस्पर्धा किस दिशा में जा रही है: ऐसे AI systems की ओर जिनका मूल्यांकन इस बात से होगा कि वे क्या पूरा कर सकते हैं, न कि सिर्फ क्या कह सकते हैं।

यह लेख The Decoder की रिपोर्टिंग पर आधारित है। मूल लेख पढ़ें.

Originally published on the-decoder.com