usama10

grpo-tax-qwen-3b-dpo

text-generationapache-2.0peft

huggingface.co/usama10/grpo-tax-qwen-3b-dpo

Updated 2026-09-21 ·Open on Hugging Face →
peft · safetensors · dpo · alignment-tax · preference-learning · lora · research · text-generation · dataset:HuggingFaceH4/ultrafeedback_binarized · base_model:Qwen/Qwen2.5-3B-Instruct · base_model:adapter:Qwen/Qwen2.5-3B-Instruct · license:apache-2.0 · region:us
The README has not been fetched yet (metadata-first ingest). The model page and the metadata below are live.
Mirrored from the Hugging Face Hub and served from the Conceptio Open Knowledge Archive. Read the original card at https://huggingface.co/usama10/grpo-tax-qwen-3b-dpo.