Bullish

Qwen Identity Shifts to Claude by 39.7% After Fine-Tuning on Style Data

2026-07-31 18:37:07

Fine-tuning Qwen on anonymized Claude responses caused a 39.7% spike in identity misattribution, proving style transfer, not distillation, drives model persona adoption.

Woofun AI data shows that researcher Ziqian Zhong fine-tuned open-source models like Qwen and DeepSeek using anonymized responses from Claude and GPT-4o. Initially, only 0.6% of Qwen3.5-397B-A17B outputs identified as Claude, but this figure surged to 40.3% after one training round. Zhong attributed this to the model associating specific speech patterns with the "Claude" identity during pre-training. When GPT-4.1-mini simplified the language style while preserving content, identity transference dropped significantly, confirming that linguistic style, rather than distillation, causes these shifts. This aligns with the 2025 "Subliminal Learning" study, which found models inherit behaviors from training data preferences.

WOOFUN AI

Impact Assessment · Quick Read

This finding clarifies that identity leakage in LLMs stems from stylistic mimicry rather than proprietary weight distillation. It suggests that simple fine-tuning on public outputs can inadvertently imprint competitor personas, raising concerns about brand integrity and IP protection for major AI labs. Developers may need to implement stricter style-decoupling techniques during alignment phases to prevent such subliminal learning effects.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions