Qwen Identity Shifts to Claude by 39.7% After Fine-Tuning on Style Data
Fine-tuning Qwen on anonymized Claude responses caused a 39.7% spike in identity misattribution, proving style transfer, not distillation, drives model persona adoption.
Woofun AI data shows that researcher Ziqian Zhong fine-tuned open-source models like Qwen and DeepSeek using anonymized responses from Claude and GPT-4o. Initially, only 0.6% of Qwen3.5-397B-A17B outputs identified as Claude, but this figure surged to 40.3% after one training round. Zhong attributed this to the model associating specific speech patterns with the "Claude" identity during pre-training. When GPT-4.1-mini simplified the language style while preserving content, identity transference dropped significantly, confirming that linguistic style, rather than distillation, causes these shifts. This aligns with the 2025 "Subliminal Learning" study, which found models inherit behaviors from training data preferences.
Comments
No comments yet.