The problem we are working on
Machine learning can build systems that nobody is able to explain, and closing that gap is what this lab works on.
Transformers, diffusion models, and the variants around them were found by experiment and then made bigger, and they work well enough that the question of why rarely gets asked. Nothing tells us they are the right design. Existing theory can analyze parts of their behavior once they exist, but it did not predict or produce any of them.
Our starting position is that computation and learning are physical processes, subject to laws about entropy and energy like anything else, and that the right designs can be calculated from those laws rather than found by trial. If that works, what comes out will not be a faster transformer. It will be a different sort of system, doing the job a different way.
There is a larger question behind that one. We have seen intelligence in two forms, living brains and the models we have built, and both are particular. If intelligence can be described mathematically without reference to either, then each is one solution to that description rather than the thing itself. Descriptions like that already exist. The best known of them cannot actually be calculated, which sounds like a dead end and is not one. It means the useful question stops being what intelligence is, and becomes which physical systems get close to it, how close, and at what cost in energy and time. Those are questions you can run an experiment on. The research program says which description we work from and where it gives out.
Start from physics, not from what is conventional
Thermodynamics, statistical mechanics, and information theory put hard limits on any learning system, whatever it is made of. Those limits are narrow enough to be useful: they rule out large classes of design before anyone writes code, which is what makes them a place to start.
Calculate the design, do not guess it
Energy landscapes, wave equations, and entropy functionals are used here as the actual mathematical objects a model is built out of. They are not illustrations of how a model behaves or analogies for explaining it afterwards.
Living intelligence is a case, not the definition
Brains are the only systems everyone agrees are intelligent, which makes them easy to mistake for what intelligence is. We read them the other way round, as one implementation that physics arrived at under constraints no designer would have chosen. That makes a brain evidence about the general case and a poor blueprint for it.
Build what the theory produces
A piece of theory is finished when it produces something that runs: a design, a procedure, or a prediction an experiment can check. Until it does, it stays on the list of things in progress.