@myq_1997 But all of this is simple pick-and-place/reorientation/open-close tasks. Hopefully this is used a good planning agent with low level policies orchestrating the complex control cause just heuristics won’t work for every task.
This Fall, I'm teaching a new "Hands-on Robot Learning" class at @JHUCompSci
A full-stack class where students will get SO-101 robot kits, build robots out, collect data, train and deploy learned policies/WAMs/agents/etc.
Materials will be posted here:
https://t.co/pvu5nxhIUU
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia?
These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach.
The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date.
The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant.
At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on “old” problems that have a long history. This history is based on assumptions about how the “vision problem” will be “solved”. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems.
So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models.
Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field.
If we want there to be a “field” of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section.
Concretely, I think papers should include a new section analogous to “Related Work” where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models.
I'm interested in your thoughts.
@stepjamUK Naive question, I don’t work with quadrupeds: did you consider RSI (Reference State Initialization) early in training, or HER (Hindsight Experience Replay) if goal-conditioned, to help sparse exploration? How would these compare to the contact-guided critic?
My group at Johns Hopkins now has a name and a website
Introducing the Brains, Bots, and Behavior Lab
https://t.co/Iokm3YxuXQ
We're always looking for creative researchers at all levels to join us!
Meet the recipients of the 2024 ACM A.M. Turing Award, Andrew G. Barto and Richard S. Sutton! They are recognized for developing the conceptual and algorithmic foundations of reinforcement learning. Please join us in congratulating the two recipients! https://t.co/GrDfgzW1fL
My paper was accepted to the 2023 IEEE International Conference on Acoustics, Speech & Signal Processing (@ieeeICASSP). Join me! https://t.co/rZzJ7MRsFF
More on cramming: CIFAR10 hyperlightspeedbench.
Train CIFAR10 to 94% in under 10 seconds on a single A100. With a single readable 600-line https://t.co/gVf4g3bzPN, bunch of nice tricks implemented within.
https://t.co/koGgN4CUKU
🪐 Introducing Galactica. A large language model for science.
Can summarize academic literature, solve math problems, generate Wiki articles, write scientific code, annotate molecules and proteins, and more.
Explore and get weights: https://t.co/jKEP8S7Yfl
Who are Sudras? Are they not Hindus? Why they have been insulted in Manusmrithi denied equality, education, employment and Temple entry. Dravidian Movement as saviour of 90% Hindus questioned and redressed these, cannot be anti-Hindus.
How can we design neural networks in a principled way? New work from @DeWeeseLab@berkeley_ai explores a theoretically-motivated paradigm for doing just that for dense networks! Read about it here: https://t.co/jPmMz9Ab9R