Lightnews — Scholar-powered news

Thaddäus Wiedemer

@thwiedemer.bsky.social

47 followers 110 following 8 posts

Intern at Google Deepmind Toronto | PhD student in ML at Max Planck Institute Tübingen and University of Tübingen.

Posts Replies Media Videos

Thaddäus Wiedemer

@thwiedemer.bsky.social

Intuitively, some tasks are easier to directly solve in the vision domain, and we also observe this in maze solving tasks. This makes me super excited about a future where generalist vision and language models could be integrated for reasoning in the real world by 'imagining' possible outcomes.

September 25, 2025 at 5:02 PM

Thaddäus Wiedemer

@thwiedemer.bsky.social

On the reasoning side, videos as 'chain-of-frames' parallel chain-of-thought in LLMs. Complex visual tasks that an image editing model like Nano Banana would have to solve in one go can be broken down into smaller steps.

September 25, 2025 at 5:02 PM

Thaddäus Wiedemer

@thwiedemer.bsky.social

Specifically, Veo 3 can perceive (segment, locacalize, detect edges, ...), model (physics, abstract relations, memory), manipulate (edit images, simulate robotics), and reason about the visual world.

Video models might well become vision foundation models.

September 25, 2025 at 5:02 PM

Thaddäus Wiedemer

@thwiedemer.bsky.social

Are we experiencing a 'GPT moment' in vision?

In our new preprint, we show that generative video models can solve a wide range of tasks across the entire vision stack without being explicitly trained for it.

🌐 video-zero-shot.github.io

1/n

September 25, 2025 at 5:02 PM

Add to Home Screen

Light up
your news

Add to Home Screen

Light upyour news

Sign in to Lightnews

Sign up to start reading

Connect Bluesky

Connect with Bluesky

Light up
your news