The question
Fixed in Note 001, before there was any result to judge against it.
Can the architectures of learning systems be derived from physical and mathematical law?
Symmetry arguments already do part of this. Bronstein et al., 2021
Parts of several architectures were reasoned rather than stumbled upon. Convolution builds in translation symmetry. Variational autoencoders come from an explicit probabilistic derivation. Diffusion models came from a thermodynamic analogy. In each case the principle fixes an objective or a symmetry, and the network carrying it is still chosen by trial. Nothing has been derived the whole way down, and that last stretch is the subject of this program.
Behind that sits a larger question. We have met intelligence in two forms, living brains and the systems we have built, and both are particular. If intelligence can be characterized mathematically without reference to either, then each is one solution to that characterization and neither is the thing itself. Definitions of that kind already exist. Legg and Hutter published one in 2007 and were explicit that it is "in no way anthropocentric", and equally explicit that its value "is not computable". A definition nobody can evaluate is a starting point rather than an answer, and it is roughly where we pick the problem up: which physical systems can approximate it, and at what cost.
The word "derived" is doing a lot of work in the question above, so here is what we mean by it, in four conditions. The assumptions are written down and you can count them. Every step can be checked by someone who did not write it. When you restrict the general structure, systems that already exist and work have to fall out of it. And somewhere the result has to disagree with current practice, on a point an experiment can settle. Something that meets fewer than four is an analogy. Analogies are useful for deciding what to try next and useless as a foundation.