Breaking News • AI • Technology • Startups • Cybersecurity • Future Tech

Photon-1: Induction Labs’ AI Learns Complex Tasks from Raw Video, No Action Labels Needed

Photon-1: Induction Labs' AI Learns Complex Tasks from Raw Video, No Action Labels Needed

Revolutionizing AI Learning from Video

At a glance, For years, training AI agents from video demonstrations typically required meticulously labeled actions for each frame. This dependency has been a significant bottleneck, limiting the scale and efficiency of AI development. However, a groundbreaking innovation from Induction Labs is challenging this paradigm with their new architecture: **Imagination Models**.

Meanwhile, Last week, Induction Labs unveiled Photon-1, a powerful test system built on this novel approach. Photon-1 learns directly from raw video, completely bypassing the need for explicit action labels during pretraining. This breakthrough promises to unlock unprecedented scalability and efficiency in how AI understands and interacts with complex digital environments.

What Exactly is an Imagination Model?

At its core, an imagination model like Photon-1 operates by predicting future frames autoregressively within a learned representation space. Instead of generating pixels during pretraining, it focuses on a ‘next-latent-token-prediction’ objective. This means the model learns to understand the underlying structure and progression of events without ever being told what specific action caused a change.

In practical terms, The critical claim here is that by simply predicting future states, the model implicitly learns to complete tasks. Induction Labs refers to this as an ‘implicit policy‘. Rather than memorizing labels for individual mouse clicks or keystrokes, Photon-1 develops a conceptual understanding of human intent and activity, allowing it to infer actions that lead to desired outcomes.

The Compression Trick Behind its Scale

A key enabler for Photon-1’s impressive capabilities and efficiency is its innovative compression technique. The architecture employs a vision encoder utilizing Finite Scalar Quantization (FSQ). This method compresses each video frame into 960 discrete tokens, with each token being an 8-dimensional vector composed of five possible values.

For example, This sophisticated encoding results in an astonishingly compact representation of about 2.2 KB per frame. Induction Labs reports over a 100x improvement in compression compared to existing OCR and multimodal-model representations, all while meticulously preserving crucial details like text, layout, and state changes within the video. To achieve this, Photon-1 uses a differential latent encoder, focusing on encoding differences between frames rather than the entire frame content.

Massive Data, Efficient Pretraining

Photon-1’s training journey began with a vast corpus of 2 billion publicly available videos. This was meticulously filtered down to approximately 2 million computer screen recordings. An internal keyframe detection model further optimized the dataset by removing redundant frames.

That said, The final dataset comprised 575 million frames, sampled at one frame per second, equating to an astounding 552 billion tokens or about 18 years of video content. Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) transformer, was pretrained from scratch for a single epoch. This intensive pretraining required approximately 30,000 H200 GPU-hours, demonstrating a remarkable computational efficiency for such a large-scale model.

From Imagination to Action

While pretraining focuses on imagination, Photon-1 can be fine-tuned to perform explicit actions. Induction Labs accomplished this by training it on fewer than 35,000 computer use trajectories, teaching it how to interpret and emit actions using special computer use tokens. During inference, Photon-1 first predicts the next frame’s state and then determines the action required to achieve that state.

Interestingly, The model further refines its capabilities through online reinforcement learning. This involves real-time rollouts on virtual machines, where outcomes are programmatically verified to provide rewards. These Linux VMs simulate diverse desktop environments (LXQt, Xfce, MATE, GNOME, and Plasma), complete with Google accounts and an internal, rate-limit-free ChatGPT clone for complex tasks.

Performance and Cost Advantages

Photon-1 demonstrates a significant advantage in both pretraining compute and inference cost compared to existing models. On an internal computer use benchmark, Induction Labs reports that Photon-1 outperforms Gemini 3.1 Flash-Lite.

  • Pretraining Compute: Photon-1 uses approximately 27x less pretraining compute (0.044 × 10²⁴ FLOPs) compared to Induction Labs’ conservative estimate for Gemini 3.1 Flash-Lite (1.200 × 10²⁴ FLOPs).
  • Weighted Inference Cost: Photon-1 also boasts a lower weighted inference cost of $0.11 per 1M tokens, significantly less than Gemini 3.1 Flash-Lite’s estimated $0.36.

However, Notably these figures include some conservative estimates and are based on an internal benchmark, meaning independent verification is not yet available.

Surprising Generalization Beyond the Desktop

Perhaps one of the most exciting aspects of Photon-1 is its ability to generalize to domains it never saw during its desktop-focused pretraining. Induction Labs conducted tests by fine-tuning Photon-1 on tasks entirely absent from its initial video corpus, comparing it against a vision encoder baseline and an LLM baseline (Ling-flash-2.0).

  • Checkers: On 20,000 tournament checkers games, Photon-1 surpassed both baselines in world simulation accuracy and move quality.
  • Billiard Physics: In 10,000 synthetically generated billiard games, Photon-1 achieved a mean absolute error of 0.47 against the ground-truth physics engine, significantly outperforming the LLM baseline (1.15) and the vision encoder baseline (1.44).

Meanwhile, Furthermore, Photon-1 showed it could pick up human priors from its pretraining videos. After reinforcement learning, it learned to effectively use the in-VM ChatGPT clone to draft documents and answer knowledge questions, mimicking how a human would interact with such a tool.

Key Innovations and Future Outlook

Photon-1 represents a major stride in AI development, demonstrating that powerful agents can learn complex behaviors and implicit policies from raw, unlabeled video data. Its innovative compression, efficient pretraining, and surprising generalization capabilities highlight a promising new direction for AI research.

In practical terms, It’s crucial to remember that Photon-1 is currently a research result – no weights, API, or license are available to the public. However, the implications for future AI systems that can learn more naturally and efficiently from the vast amounts of video data available are immense.

Expert Perspective

A practical read on Photon-1 starts with photon. That is where the earliest effects are likely to show up if this development keeps building.

What happens next will come down to adoption speed, policy response, and execution quality. That combination could make Photon-1 a meaningful reference point across induction.

For decision-makers, the useful lens is not the headline alone but how labs changes priorities once organizations have to respond.

Frequently Asked Questions

Why is Photon-1 important?

Revolutionizing AI Learning from VideoAt a glance, For years, training AI agents from video demonstrations typically required meticulously labeled actions for each frame.

What impact could Photon-1 have?

This dependency has been a significant bottleneck, limiting the scale and efficiency of AI development.

What should readers watch next with Photon-1?

However, a groundbreaking innovation from Induction Labs is challenging this paradigm with their new architecture: **Imagination Models**.Meanwhile, Last week, Induction Labs unveiled Photon-1, a powerful test system built on this novel approach.

How does this relate to photon?

It connects because the article frames photon as one of the clearest areas where the topic may be felt in practice.

Source: https://www.marktechpost.com/2026/07/26/induction-labs-photon-1-simulates-desktops-plays-checkers-and-models-billiard-physics-from-one-pretraining-run/

Share this article

Subscribe

By pressing the Subscribe button, you confirm that you have read our Privacy Policy.

Latest News

More Articles