genpark-sac-maximum-entropy-rl-evaluator-skill
mcp
Uyari
Health Uyari
- No license — Repository has no license file
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Gecti
- Code scan — Scanned 6 files during light audit, no dangerous patterns found
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Soft Actor-Critic (SAC) maximum entropy policy value and soft Bellman target evaluator
README.md
genpark-sac-maximum-entropy-rl-evaluator-skill
Agent Skill implementing the Soft Actor-Critic (SAC) Maximum Entropy Reinforcement Learning Formulation, incentivizing policy exploration through Shannon entropy regularization.
Architectural Overview
flowchart TD
QValues["Action-Value Q(s, a)"] & Probs["Action Probabilities pi(a|s)"] --> Entropy["Compute Entropy H(pi) = - sum pi log pi"]
QValues & Probs & Entropy --> SoftVal["Soft State Value: V(s) = sum pi(a|s)[Q(s, a) - alpha * log pi(a|s)]"]
SoftVal & Reward["Reward r"] --> SoftTarget["Soft Bellman Target: y = r + gamma * V(s')"]
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi