coming from physics, this is something i do not particularly like about ML research, especially LLMs: most research ideas/approaches are just trial and error or a permutation and combination of architecures/training objectives etc. researchers don't know why something works or how, nobody does. it's not their fault. but it's just something i find irritating. like a new permutation of architectural designs can completely overthrow a previous SOTA model without any massive changes. there are too many degrees of freedom and the design space is HUGE.
adapting to a research field that operates on empiricism, while coming from one that operates on explanatory frameworks gives me massive unrest and leaves me in confusion. theories have longevity in physics/math. while in ML, paradigms change weekly.