genpark-dpo-preference-pair-dataset-generator-skill

mcp
Security Audit
Warn
Health Warn
  • No license — Repository has no license file
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 7 GitHub stars
Code Pass
  • Code scan — Scanned 6 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Direct Preference Optimization (DPO) and RLHF dataset constructor pairing prompt, chosen, and rejected responses

README.md

genpark-dpo-preference-pair-dataset-generator-skill

High-integrity dataset generator and validator for Direct Preference Optimization (DPO) and RLHF alignment workflows.

Architecture

flowchart LR
    Prompt[User Prompt] --> Generator[DPOPreferenceGenerator]
    Chosen[Chosen Response] --> Generator
    Rejected[Rejected Response] --> Generator
    Generator --> Validator[Validation & Margin Engine]
    Validator --> Export[JSONL DPO Dataset]

Features

  • Strict Quality Constraints: Validates text uniqueness, non-empty criteria, and score margins.
  • JSONL Serialization: Formats clean standard DPO dataset lines for direct consumption by Hugging Face TRL or Axolotl.
  • Standard Library Only: No external dependencies.

Reviews (0)

No results found