మరింత విస్తృత control modelతో Google తన robotics stackను విస్తరిస్తోంది
Google DeepMind Gemini Robotics 2ను పరిచయం చేసింది, ఇది ఒక కొత్త vision-language-action model; కంపెనీ ప్రకారం ఇది చాలా భిన్నమైన robotsకు ఒక common control layerగా పనిచేయగలదు. మూల పదార్థంలో వివరించిన ప్రకటన ప్రకారం, ఈ model tabletop arms నుంచి full-body humanoid robots వరకు ఉన్న systemsలో పని చేయడానికి ఉద్దేశించబడింది, physical worldలో పనిచేయగల softwareగా పెద్ద multimodal modelsను మార్చాలనే Google ప్రయత్నాన్ని ఇది ముందుకు తీసుకెళ్తోంది.
ఈ release ముఖ్యమైనది, ఎందుకంటే robotics developers చాలా కాలంగా fragmentation సమస్యను ఎదుర్కొంటున్నారు. ఒక machine type లేదా taskలో బాగా పనిచేసే models, వేరే body plan, sensor setup, లేదా environmentను నిర్వహించడానికి ముందు తరచూ గణనీయమైన adaptation అవసరం అవుతుంది. DeepMind Gemini Robotics 2ను మరింత general foundation దిశగా ఒక అడుగుగా చూపిస్తోంది: imagesను అర్థం చేసుకోగల, languageను process చేయగల, మరియు ఆ inputsను విస్తృత hardware శ్రేణిలో actionsగా మార్చగల ఒక model family.
ఈ positioning technical claimకే పరిమితం కాకుండా కూడా ముఖ్యం. ప్రాక్టికల్ పరంగా, robots కోసం ఒక “intelligence layer” అంటే core perception, reasoning, మరియు control software అనేక platformsలో తిరిగి ఉపయోగించగల marketను సూచిస్తుంది, hardware makers మాత్రం mechanics, reliability, cost, మరియు deployment ద్వారా తేడా చూపిస్తారు. ఈ approach వాస్తవ ప్రపంచ పరీక్షల్లో నిలబడితే, logistics, industrial handling, research, మరియు service applications కోసం adaptive robots నిర్మించడంలో అడ్డంకిని తగ్గించగలదు.
కొత్త model ఏమి చేయగలదని DeepMind చెబుతోంది
మూల పాఠ్యం Gemini Robotics 2ను ఇప్పటి వరకు DeepMind యొక్క అత్యంత advanced vision-language-action modelగా వివరిస్తోంది. Vision-language-action systems దృశ్య అవగాహన, సహజ భాషా వ్యాఖ్యానం, మరియు motor controlను కలిపి, robot తాను ఏమి చూస్తుందో, దానికి ఏమి చేయమని చెప్పారో అనుసంధానించి, తర్వాత physical responseను అమలు చేయడానికి వీలు కల్పిస్తాయి.
కొత్త model full-body movementను నిర్వహించగలదని, fine motor tasks చేయగలదని, మరియు multiple robotsను coordinate చేయగలదని DeepMind చెబుతోంది. ఇవి మూడు చాలా భిన్నమైన అవసరాలు. Full-body motionకు విస్తృత spatial planning మరియు balance-related control అవసరం. Fine motor work మరింత ఖచ్చితమైన manipulationపై ఆధారపడుతుంది. Multi-robot coordination timing మరియు shared-task complexityను తీసుకువస్తుంది. ఈ సామర్థ్యాలను ఒకే platformలో కుదించడం Gemini Robotics 2 కేవలం incremental upgrade కాదు, మరింత విస్తృత control architecture అని కంపెనీ వాదనలో కేంద్ర భాగం.
మూలం developers early access కోసం waitlist ద్వారా apply చేయవచ్చని కూడా చెబుతోంది. ఇది advanced AI systemsలో సాధారణ rollout patternను సూచిస్తుంది: ముందు public announcement, తర్వాత controlled access, ఆపై performance మరియు safety సరైనవిగా తేలితే wider deployment. Developersకు near-term implication immediate production deployment కంటే evaluation, prototyping, మరియు model ప్రస్తుత robotics stacksలో ఎక్కడ సరిపోతుందో తెలుసుకోవడంపై ఎక్కువగా ఉంటుంది.
రెండవ model embodied reasoningపై దృష్టి పెట్టింది
Gemini Robotics 2తో పాటు, Google DeepMind Gemini Robotics ER 2ను కూడా పరిచయం చేసింది. “ER” అనే లేబుల్ embodied reasoningను సూచిస్తుంది, దీనిని మూల పాఠ్యం physical worldను అర్థం చేసుకుని, ఆ అవగాహన ఆధారంగా ఎలాంటి actions తీసుకోవాలో నిర్ణయించడంగా వివరిస్తోంది. మరో మాటలో చెప్పాలంటే, ఈ model low-level execution కంటే robotic systems కోసం higher-level interpretation మరియు decision supportపై ఎక్కువ దృష్టి పెట్టింది.
ER 2, Aprilలో విడుదలైన Gemini Robotics ER 1.6ను భర్తీ చేస్తుందని DeepMind చెబుతోంది. ఈ వేగవంతమైన succession, general AI capability మరియు robotics-specific tuning మిశ్రమంపై model usefulness ఆధారపడే రంగంలో కంపెనీ త్వరగా iterate చేస్తోందని సూచిస్తోంది. ముఖ్యమైన distribution వివరమేమిటంటే ER 2 Google AI Studioలో అందుబాటులో ఉంది, దీంతో developersకు stack యొక్క reasoning sideను పరీక్షించడానికి మరింత ప్రత్యక్ష మార్గం లభిస్తుంది; broader Robotics 2 control model మాత్రం early access ద్వారా నియంత్రించబడుతోంది.
ఒక general control model మరియు ఒక reasoning modelగా విభజించడం robotics AIలో విస్తృత design trendను కూడా ప్రతిబింబిస్తుంది. Companies increasingly “robot ఏమి చేయాలి?” మరియు “robot దాన్ని physicalగా ఎలా చేయాలి?” అనే సమస్యలను వేరు చేస్తున్నాయి. ఒక reasoning system tasksను విడదీయడంలో, scenesను interpret చేయడంలో, మరియు sequencesను plan చేయడంలో సహాయపడగలదు, అయితే ఒక control model ఆ plansను movementగా మార్చగలదు. DeepMind product structure ఆ విభజనకు అనుగుణంగా కనిపిస్తోంది.
Launch ఎందుకు ముఖ్యమైనది
ఈ announcement తదుపరి తరం robots కోసం software foundationను నిర్వచించడానికి పెరుగుతున్న పోటీలో మరో అంశాన్ని జోడిస్తోంది. పరిశ్రమ కఠినంగా scripted automation నుంచి, తక్కువ predictability ఉన్న settingsకు అనుగుణంగా మారగల systems వైపు కదులుతోంది. Warehouses, factories, labs, మరియు చివరికి public-facing environments అన్ని edge caseలకు exhaustive manual programming అవసరం లేకుండా variationను నిర్వహించగల robotsను కోరుకుంటాయి.
DeepMind framing, multimodal foundation modelsనే ఆ adaptabilityకి మార్గంగా చూస్తోందని సూచిస్తోంది. ఒక robot language instructionsను అర్థం చేసుకోగలిగితే, visual sceneను parse చేయగలిగితే, మరియు tasks across actionsను generalize చేయగలిగితే, developers తక్కువ custom codeతో మరింత flexible productsను నిర్మించగలరు. ఇది safety, latency, calibration, hardware integration, మరియు evaluation చుట్టూ ఉన్న కఠిన engineering workను తొలగించదు. కానీ stackలో అత్యంత కఠిన సమస్యలు ఎక్కడ ఉంటాయో మార్చగలదు.
Googleకు ఇక్కడ ఒక strategic point కూడా ఉంది. Software వైపు generalize చేయడంలో ఇబ్బంది ఉండటంతో robotics కొన్ని సార్లు వేగంగా పెరిగి, ఆపై నిలిచిపోయింది. Roboticsను కంపెనీ flagship AI model ecosystemతో మరింత నేరుగా అనుసంధానించడం ద్వారా, Google DeepMind general multimodal AIలోని పురోగతిని embodied systemsలో పురోగతిగా మార్చవచ్చని వాదిస్తోంది. Roboticsను ప్రధానంగా వేరే research trackగా చూడటానికి ఇది మరింత బలమైన commercial story.
తరువాత ఏమి గమనించాలి
- Early-access developers నిర్దిష్ట robotsకు మాత్రమే పరిమితమైన demos కంటే బలమైన cross-platform performanceను నివేదిస్తున్నారా.
- విభిన్న embodiments మరియు safety constraintsకు Gemini Robotics 2ను అనుకూలీకరించడానికి ఇంకా ఎంత integration work అవసరమో.
- Google AI Studioలో ER 2, DeepMind immediate ecosystem బయట ఉన్న robotics developersకు practical planning toolగా మారుతుందా.
- General-purpose robot software platforms నిర్మించడానికి పోటీ వేగం పెరుగుతున్నప్పుడు rivals ఎలా స్పందిస్తారు.
ప్రస్తుతం, ఈ launchను ప్రతి చోటా విస్తృత సామర్థ్యం ఉన్న robots సిద్ధంగా ఉన్నాయనే రుజువుగా కాకుండా ఒక ముఖ్యమైన product మరియు platform signalగా చదవడం మంచిది. కానీ ఇది పెద్ద AI labsలో ఒకటి market ఏ దిశగా వెళ్తోందని నమ్ముతోందో చూపిస్తుంది: అనేక రకాల machines పైన నిలబడగల, తిరిగి ఉపయోగించగల, multimodal intelligence layers వైపు, వాటిని వాస్తవ ప్రపంచంలో మరింత adaptableగా మార్చే దిశగా.
ఈ article The Decoder reportపై ఆధారపడింది. మూల article చదవండి.
Originally published on the-decoder.com


