Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
Standard post-training pipelines like SFT and RL can inadvertently degrade specific values instilled during pre-training.
Researchers analyzed how different post-training domains affect the retention of compassion values in a Llama 3.1 8B model. They found that applying helpfulness-oriented SFT or RL can erode previously learned values, suggesting that the choice of post-training data significantly impacts model alignment and value stability.