What's in the Gallery
What the project does is plain but painstaking: draw the architectures of the major models—attention-mechanism variants, the placement of normalization, MoE routing schemes, positional-encoding choices—in a unified legend, side by side and comparable. When reading papers, these details are scattered across text and formulas, and only by looking at them aligned can you spot the pattern: the differences among architectures actually concentrate on a handful of design points, and most components have already converged into industry consensus. For learners, this is more efficient than grinding through ten papers; for practitioners, it's a rare "archaeological map."
The Current State of Convergence and Divergence
Read the whole gallery through and you get an interesting judgment: after the Transformer, at the architecture level there are few revolutions and much evolution, and the industry's competitive center of gravity long ago shifted to data, training methods, and engineering efficiency—the architecture diagrams look increasingly alike while the models pull further apart, showing the decisive factor isn't on the diagram. But the other half of the gallery's value is precisely in recording the sprouts of divergence: "non-mainstream" branches like state-space models and hybrid architectures are also catalogued, and should a paradigm shift arrive someday, looking back at this gallery is the historical scene. Worth a spot in your bookmarks—it's still being updated.
via: Hacker News