genpark-sac-maximum-entropy-rl-evaluator-skill
mcp
Warn
Health Warn
- No license — Repository has no license file
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Pass
- Code scan — Scanned 6 files during light audit, no dangerous patterns found
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
Soft Actor-Critic (SAC) maximum entropy policy value and soft Bellman target evaluator
README.md
genpark-sac-maximum-entropy-rl-evaluator-skill
Agent Skill implementing the Soft Actor-Critic (SAC) Maximum Entropy Reinforcement Learning Formulation, incentivizing policy exploration through Shannon entropy regularization.
Architectural Overview
flowchart TD
QValues["Action-Value Q(s, a)"] & Probs["Action Probabilities pi(a|s)"] --> Entropy["Compute Entropy H(pi) = - sum pi log pi"]
QValues & Probs & Entropy --> SoftVal["Soft State Value: V(s) = sum pi(a|s)[Q(s, a) - alpha * log pi(a|s)]"]
SoftVal & Reward["Reward r"] --> SoftTarget["Soft Bellman Target: y = r + gamma * V(s')"]
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found