genpark-length-bias-penalty-reward-calibrator-skill
mcp
Uyari
Health Uyari
- No license — Repository has no license file
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Gecti
- Code scan — Scanned 6 files during light audit, no dangerous patterns found
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Verbosity penalty and length bias calibrator normalizing raw model rewards against response token lengths
README.md
genpark-length-bias-penalty-reward-calibrator-skill
Verbosity penalty calibrator discounting inflated reward signals associated with excessively long model outputs.
Architecture
flowchart LR
Length[Token Length] --> Excess[max(0, Length - Target)]
Excess --> Penalty[Penalty = alpha * Excess]
Raw[Raw Model Reward] --> Discount[Calibrated Reward = Raw - Penalty]
Penalty --> Discount
Features
- Linear and Exponential Decay: Flexible penalty profiles.
- Zero Dependencies: 100% Python Standard Library.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi