Thursday, 15 May 2025
  • My Feed
  • My Interests
  • My Saves
  • History
  • Blog
Subscribe
Capernaum
  • Finance
    • Cryptocurrency
    • Stock Market
    • Real Estate
  • Lifestyle
    • Travel
    • Fashion
    • Cook
  • Technology
    • AI
    • Data Science
    • Machine Learning
  • Health
    HealthShow More
    Foods That Disrupt Our Microbiome
    Foods That Disrupt Our Microbiome

    Eating a diet filled with animal products can disrupt our microbiome faster…

    By capernaum
    Skincare as You Age Infographic
    Skincare as You Age Infographic

    When I dove into the scientific research for my book How Not…

    By capernaum
    Treating Fatty Liver Disease with Diet 
    Treating Fatty Liver Disease with Diet 

    What are the three sources of liver fat in fatty liver disease,…

    By capernaum
    Bird Flu: Emergence, Dangers, and Preventive Measures

    In the United States in January 2025 alone, approximately 20 million commercially-raised…

    By capernaum
    Inhospitable Hospital Food 
    Inhospitable Hospital Food 

    What do hospitals have to say for themselves about serving meals that…

    By capernaum
  • Sport
  • 🔥
  • Cryptocurrency
  • Data Science
  • Travel
  • Real Estate
  • AI
  • Technology
  • Machine Learning
  • Stock Market
  • Finance
  • Fashion
Font ResizerAa
CapernaumCapernaum
  • My Saves
  • My Interests
  • My Feed
  • History
  • Travel
  • Health
  • Technology
Search
  • Pages
    • Home
    • Blog Index
    • Contact Us
    • Search Page
    • 404 Page
  • Personalized
    • My Feed
    • My Saves
    • My Interests
    • History
  • Categories
    • Technology
    • Travel
    • Health
Have an existing account? Sign In
Follow US
© 2022 Foxiz News Network. Ruby Design Company. All Rights Reserved.
Home » Blog » This AI Paper from NVIDIA Introduces Cosmos-Reason1: A Multimodal Model for Physical Common Sense and Embodied Reasoning
AIMachine LearningTechnology

This AI Paper from NVIDIA Introduces Cosmos-Reason1: A Multimodal Model for Physical Common Sense and Embodied Reasoning

capernaum
Last updated: 2025-03-25 04:58
capernaum
Share
This AI Paper from NVIDIA Introduces Cosmos-Reason1: A Multimodal Model for Physical Common Sense and Embodied Reasoning
SHARE

Artificial intelligence systems designed for physical settings require more than just perceptual abilities—they must also reason about objects, actions, and consequences in dynamic, real-world environments. These systems must understand spatial arrangements, cause-and-effect relationships, and the progression of events over time. In applications like robotics, self-driving vehicles, or assistive technologies, AI must comprehend its surroundings’ physical constraints and affordances to make intelligent and safe decisions. This fusion of perception with structured reasoning about physical dynamics forms the backbone of Physical AI.

A core issue for such systems is their inability to conclude physical environments using integrated visual and contextual information. Although vision-language models have made significant progress, they still struggle to determine whether a task has been completed, what action should follow next, or whether a proposed action is feasible. The gap between perception and decision-making becomes especially critical when AI needs to operate independently and interpret tasks from complex visual scenarios. These systems remain unreliable in high-stakes or fast-changing environments without mechanisms to verify their reasoning.

Existing models such as LLaVA, GPT-4o, and Gemini 2.0 Flash are proficient in handling text and visual data but underperform physically grounded reasoning. Tasks like identifying temporal order, spatial continuity, or object permanence are rarely handled effectively. Popular benchmarks often fail to evaluate such scenarios, offering limited insight into a model’s ability to reason about physical events or agent actions. Moreover, current systems usually rely on textual cues rather than making decisions based on visual evidence, leading to inconsistent or incorrect conclusions when applied to the physical world.

Researchers from NVIDIA introduced Cosmos-Reason1, a family of vision-language models developed specifically for reasoning about physical environments. These models were released in two sizes: 8 billion and 56 billion parameters. The models were built with a structured approach that included defining ontologies for physical common sense, constructing specialized training data, and designing a comprehensive suite of evaluation benchmarks. These benchmarks test capabilities such as action prediction, task verification, and judgment of physical feasibility. The research team developed datasets including BridgeData V2, RoboVQA, RoboFail, AgiBot, HoloAssist, and AV to rigorously evaluate the models.

Cosmos-Reason1 uses a hybrid Mamba-MLP-Transformer architecture that integrates both vision and language components. The training process was conducted in multiple phases. Initially, a vision encoder and language model were pretrained and fine-tuned using general supervised data. Then, a physical AI-specific supervised fine-tuning (SFT) phase introduced datasets focused on space, time, and object interactions. The final reinforcement learning (RL) phase applied rule-based rewards to improve performance in areas like arrow of time detection, spatial puzzles, and object permanence. The RL setup used a modular framework that leveraged distributed computing to scale training efficiently. The model responses were structured using tags, allowing reward systems to evaluate both correctness and reasoning structure. Each question had up to nine model-generated responses, and RL training continued for 500 iterations using a global batch size of 128 questions.

Evaluation of Cosmos-Reason1 showed a substantial performance increase compared to other models. In the physical common sense benchmark, Cosmos-Reason1-56B achieved an average accuracy of 60.2%, outperforming OpenAI o1, which scored 59.9%. The 8B variant also improved, reaching 52.3%. Cosmos-Reason1-56B scored an average of 63.7% for embodied reasoning tasks, up from a 53.5% baseline. Benchmarks like RoboVQA and HoloAssist showed strong gains, with the 56B model scoring 80.0% and 57.8%, respectively. Cosmos-Reason1-8B improved to 68.7% on intuitive physics tasks, showing strong gains in object permanence and spatial puzzle reasoning. However, the model faced challenges on datasets like RoboFail due to a lack of sufficiently diverse training examples.

In conclusion, this research introduces a targeted and layered strategy to advance AI systems that reason about physical interactions. The researchers at NVIDIA created a scalable training method combined with a comprehensive evaluation to tackle long-standing gaps in embodied reasoning. Cosmos-Reason1 demonstrates how structured fine-tuning and reinforcement learning can build AI systems more aligned with real-world physical logic and agent behavior.


Check out the Paper and GitHub Page. All credit for this research goes to the researchers of this project. Also, feel free to follow us on Twitter and don’t forget to join our 85k+ ML SubReddit.

The post This AI Paper from NVIDIA Introduces Cosmos-Reason1: A Multimodal Model for Physical Common Sense and Embodied Reasoning appeared first on MarkTechPost.

Share This Article
Twitter Email Copy Link Print
Previous Article United & Chase Refresh Cobranded Card Portfolio With Higher fees & Structured Credits United & Chase Refresh Cobranded Card Portfolio With Higher fees & Structured Credits
Next Article Mt. Gox Moves 11,500 Bitcoins Again, Major BTC Price Volatility Ahead? Mt. Gox Moves 11,500 Bitcoins Again, Major BTC Price Volatility Ahead?
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Your Trusted Source for Accurate and Timely Updates!

Our commitment to accuracy, impartiality, and delivering breaking news as it happens has earned us the trust of a vast audience. Using RSS feeds, we aggregate news from trusted sources to ensure real-time updates on the latest events and trends. Stay ahead with timely, curated information designed to keep you informed and engaged.
TwitterFollow
TelegramFollow
LinkedInFollow
- Advertisement -
Ad imageAd image

You Might Also Like

Gateless integrates with Fannie Mae’s income calculator

By capernaum
Exclusive Talk: Joey Conway of NVIDIA on Llama Nemotron Ultra and Open Source Models
AI

Exclusive Talk: Joey Conway of NVIDIA on Llama Nemotron Ultra and Open Source Models

By capernaum

Building AI Agents? A2A vs. MCP Explained Simply

By capernaum

Virtuo homebuyer concierge platform launches in Texas

By capernaum
Capernaum
Facebook Twitter Youtube Rss Medium

Capernaum :  Your instant connection to breaking news & stories . Stay informed with real-time coverage across  AI ,Data Science , Finance, Fashion , Travel, Health. Your trusted source for 24/7 insights and updates.

© Capernaum 2024. All Rights Reserved.

CapernaumCapernaum
Welcome Back!

Sign in to your account

Lost your password?