Researchers spliced two unrelated language models together at the layer level and got a working, if worse, hybrid.
The project, called NinaXander, takes a frozen RWKV-4-Raven-7B model (a recurrent architecture) and a frozen Tulu-Pythia-6.9b model (a Transformer) and connects them with a single trained adapter that translates between their internal representations. Run the first few layers of one model, convert the output once, then finish with the other model's remaining layers. One trained adapter supports multiple splice points without retraining. The best configuration ran Pythia's first 5 layers into RWKV's last 27, cutting the Transformer's key-value cache by 84.4% with accuracy statistically indistinguishable from RWKV alone.
That cache reduction is the one genuinely useful number here, since KV cache is a real memory bottleneck for Transformer inference at scale. But it comes with a catch: no spliced model matched Pythia's own multiple-choice accuracy, and performance on WikiText, a dataset outside the adapter's training domain, fell off a cliff. The researchers also found their representation alignment worked cleanly in only one setup, where both models happened to share a tokenizer, depth, and hidden width - which is a long way from evidence that different model families share some universal semantic space.
This reads less like a breakthrough and more like a careful negative result with one bright spot. Frankenstein-ing models together has been a recurring fantasy in open-source AI circles, usually oversold; this paper is useful precisely because it quantifies where it breaks rather than hyping the splice that happened to work.