genpark-dpo-preference-pair-dataset-generator-skill
mcp
Warn
Health Warn
- No license — Repository has no license file
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Pass
- Code scan — Scanned 6 files during light audit, no dangerous patterns found
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
Direct Preference Optimization (DPO) and RLHF dataset constructor pairing prompt, chosen, and rejected responses
README.md
genpark-dpo-preference-pair-dataset-generator-skill
High-integrity dataset generator and validator for Direct Preference Optimization (DPO) and RLHF alignment workflows.
Architecture
flowchart LR
Prompt[User Prompt] --> Generator[DPOPreferenceGenerator]
Chosen[Chosen Response] --> Generator
Rejected[Rejected Response] --> Generator
Generator --> Validator[Validation & Margin Engine]
Validator --> Export[JSONL DPO Dataset]
Features
- Strict Quality Constraints: Validates text uniqueness, non-empty criteria, and score margins.
- JSONL Serialization: Formats clean standard DPO dataset lines for direct consumption by Hugging Face TRL or Axolotl.
- Standard Library Only: No external dependencies.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found