Direct Preference Optimization Beyond Chatbots
Direct Preference Optimization Beyond Chatbots Direct Preference Optimization Beyond Chatbots Team Article Published June 3, 2026 Upvote 12 Erick Lachmann ErickvL Dharma-AI Gabriel Pimenta de Freitas Cardoso GabrielPimenta99 Dharma-AI Francisco de Almeida Rocha Alves falves9101 Dharma-AI Using Rejection Pairs From Your Model's Own Failures In April, we released DharmaOCR, our specialized structured OCR model ( available on Hugging Face ) along with a paper detailing the methodology behind it and a benchmark demonstrating its superior quality and cost efficiency. The paper benchmarked leading vision-language model families - both open-source and commercial - on a structured document extraction task: OCR on Brazilian Portuguese text.
This Research is relevant to the technology intelligence record because it involves Hugging Face, Cohere, Qwen. The source article should remain the factual reference for follow-up coverage.
- Direct Preference Optimization Beyond Chatbots Team Article Published June 3, 2026 Upvote 12 Erick Lachmann ErickvL Dharma-AI Gabriel Pimenta de Freitas Cardoso GabrielPimenta99 Dharma-AI Francisco de Almeida Rocha Alves falves9101 Dharma-AI Using Rejection Pairs From Your Model's Own Failures In April, we released DharmaOCR, our specialized structured OCR model ( available on Hugging Face ) along with a paper detailing the methodology behind it and a benchmark demonstrating its superior quality and cost efficiency.
- The paper benchmarked leading vision-language model families - both open-source and commercial - on a structured document extraction task: OCR on Brazilian Portuguese text.
- Among the reported metrics was text degeneration rate: the frequency with which a model produces a repetition loop instead of a transcription.
- Across the tested open-source families, vanilla degeneration rates ranged from below 1% to above 33%.
- Supervised fine-tuning reduced those rates for most models - but rarely to production-acceptable levels.
- The pattern points to a structural limitation: SFT optimizes for correct outputs, but does not explicitly penalize degeneration.