genpark-dpo-preference-pair-dataset-generator-skill
mcp
Uyari
Health Uyari
- No license — Repository has no license file
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Gecti
- Code scan — Scanned 6 files during light audit, no dangerous patterns found
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Direct Preference Optimization (DPO) and RLHF dataset constructor pairing prompt, chosen, and rejected responses
README.md
genpark-dpo-preference-pair-dataset-generator-skill
High-integrity dataset generator and validator for Direct Preference Optimization (DPO) and RLHF alignment workflows.
Architecture
flowchart LR
Prompt[User Prompt] --> Generator[DPOPreferenceGenerator]
Chosen[Chosen Response] --> Generator
Rejected[Rejected Response] --> Generator
Generator --> Validator[Validation & Margin Engine]
Validator --> Export[JSONL DPO Dataset]
Features
- Strict Quality Constraints: Validates text uniqueness, non-empty criteria, and score margins.
- JSONL Serialization: Formats clean standard DPO dataset lines for direct consumption by Hugging Face TRL or Axolotl.
- Standard Library Only: No external dependencies.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi