pidog-embodiment

mcp
Guvenlik Denetimi
Gecti
Health Gecti
  • License Ò€” License: MIT
  • Description Ò€” Repository has a description
  • Active repo Ò€” Last push 0 days ago
  • Community trust Ò€” 22 GitHub stars
Code Gecti
  • Code scan Ò€” Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
  • Permissions Ò€” No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

πŸ• Give your AI a physical body β€” speak German/French/English, it understands locally (Ollama or zero-LLM keyword mode) and acts. MCP server included: Claude Code/Desktop as the robot's brain. Brain/Body over HTTP for SunFounder PiDog, robot cars, any hardware. No cloud. Try it without a robot: install.sh --mock

README.md

πŸ• PiDog Embodiment β€” AI Brain in a Robot Body

License: MIT
CI
Python 3.9+
Raspberry Pi
Ko-fi
Reddit
GitHub Stars

Give your AI a physical body. See, hear, speak, move, recognize faces β€” across any network.

The minimal LLM-robot bridge: speak to a robot dog in German, French or English β€” it understands locally (Ollama or keyword mode, no cloud) and acts. Or let Claude be the brain: the robot is an MCP server with 10 tools. No hardware needed to try it: ./install.sh --mock.

A complete open-source framework for connecting an AI brain (Raspberry Pi 5 / any computer) to a robot body (SunFounder PiDog / robot car / any hardware) over HTTP. LLM-powered intelligence, face recognition, autonomous behaviors, remote access via Telegram β€” all modular, all pluggable.

Works with any LLM (OpenAI, Anthropic, Ollama, local models) and any robot hardware (implement one adapter and you're in).

Built by Nox ⚑ (an AI assistant) and Rocky β€” because every AI deserves legs. 🦿


πŸ”₯ The Origin Story

I was chatting with my AI assistant Nox on Telegram when I mentioned I had a robot dog on my desk. Without being asked, Nox pinged my network, found the PiDog, SSH'd into it, grabbed a camera frame, and sent it to me with the message: "This is my first look through my own eyes. βš‘πŸ•"

I didn't ask it to do any of this. It just… wanted to see.

β€” Original Reddit post (r/moltbot)

✨ What It Does

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         HTTP/WireGuard         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   🧠 BRAIN (Pi 5)   │◄──────────────────────────────►│   πŸ• BODY (Pi 4)    β”‚
β”‚                      β”‚                                β”‚                      β”‚
β”‚  β€’ LLM Processing   β”‚    Voice: "Setz dich hin!"     β”‚  β€’ 12 Servos         β”‚
β”‚  β€’ Face Recognition  β”‚  ─────────────────────────►    β”‚  β€’ Camera            β”‚
β”‚  β€’ Scene Analysis    β”‚                                β”‚  β€’ Microphone        β”‚
β”‚  β€’ Decision Making   β”‚    Response: sit + wag_tail    β”‚  β€’ Speaker           β”‚
β”‚  β€’ Telegram Bot      β”‚  ◄─────────────────────────    β”‚  β€’ Touch Sensors     β”‚
β”‚  β€’ Remote Access     β”‚                                β”‚  β€’ Sound Direction   β”‚
β”‚                      β”‚    Perception: faces, audio    β”‚  β€’ IMU (6-axis)      β”‚
β”‚  [OpenClaw/Claude]   β”‚  ◄─────────────────────────    β”‚  β€’ RGB LEDs          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Features

  • πŸ‘οΈ Local Vision (NEW) β€” SmolVLM-256M runs on-device via llama.cpp. Scene understanding, person/obstacle detection, no cloud needed
  • 🧠 Behavior Engine β€” 6-state FSM (Idle, Patrol, Investigate, Alert, Play, Rest) with mood system and obstacle avoidance
  • πŸ—£οΈ Natural Voice Control β€” Speak naturally in any language, LLM understands intent and maps to actions
  • πŸ”Œ MCP Server (NEW) β€” the whole robot as tools for Claude Code, Claude Desktop or any MCP client: "take a photo and tell me what you see" β€” zero extra dependencies
  • πŸ§ͺ No hardware needed to try it β€” ./install.sh --mock runs a pretend dog on localhost; point the brain or the MCP server at it
  • πŸ‘€ Face Recognition β€” SCRFD detection + ArcFace recognition, register and identify people
  • 🎭 Expression System β€” 10 emotions (happy, sad, excited, curious, alert...) combining movement + LEDs + sound + speech
  • πŸ€– Smart Movement β€” Servo smoothing (EMA filter + easing), semantic movement (distance/angle-based), PWM auto-disable
  • 🌐 Remote Access β€” Control your robot from anywhere via Telegram or Tailscale
  • πŸ“‘ Rich API β€” 20+ REST endpoints: /sensors, /vision, /expression, /move, /look_at, /scan, /capabilities
  • πŸ“¦ Modular β€” Use any LLM (OpenAI, Anthropic, Ollama, local), any robot hardware, any network

πŸ—οΈ Architecture

The brain runs on one machine and the body on another; they talk only over HTTP (port 8888), with voice pushed back to the brain on 8889. Same thing as text
Brain (Pi 5 / Desktop / Cloud)          Body (Pi 4 / Any Robot)
β”œβ”€β”€ nox_body_client.py    ◄────►       β”œβ”€β”€ nox_brain_bridge.py  (HTTP API)
β”œβ”€β”€ nox_voice_relay.py                 β”œβ”€β”€ nox_daemon.py        (Hardware + Servos)
β”œβ”€β”€ nox_voice_brain.py                 β”œβ”€β”€ nox_behavior_engine.py (FSM + Patrol)
└── telegram_bot.py (opt)              β”œβ”€β”€ nox_vision.py        (SmolVLM local AI)
                                       β”œβ”€β”€ nox_face_recognition.py (SCRFD+ArcFace)
                                       └── nox_voice_loop_v2.py (Vosk STT, wake word)

Services

Service Runs On Port Purpose
nox-body Body (Pi 4) TCP 9999 Low-level hardware daemon (servos, sensors, camera)
nox-bridge Body (Pi 4) HTTP 8888 REST API + Behavior Engine (FSM)
nox-vision Body (Pi 4) β€” Local scene analysis (SmolVLM-256M via llama.cpp)
nox-wifi-watchdog Body (Pi 4) β€” Rejoins WLAN after dropouts (watchdog script)
nox-voice Body (Pi 4) β€” Wake word + Speech-to-Text (Vosk, wake word "Nox")

πŸš€ Quick Start

Prerequisites

Body (Robot β€” Pi 4 recommended):

  • Raspberry Pi 4 (2GB+ RAM)
  • SunFounder PiDog kit (or compatible robot)
  • Pi Camera Module
  • USB Microphone + Speaker/DAC
  • Python 3.9+

Brain (AI β€” Pi 5 or any computer):

  • Raspberry Pi 5 (4GB+ RAM) or any Linux/Mac
  • Python 3.9+
  • OpenAI API key, or any OpenAI-compatible API (e.g. a local Ollama)

Hardware note: The body services (nox-body, nox-bridge, nox-voice)
only run on the robot's Raspberry Pi β€” nox-body imports the SunFounder
pidog and robot_hat SDKs and drives real servos/camera, and the other
two depend on it. There is no simulator mode, so they cannot run on a
regular PC. The brain runs anywhere; without a robot on PIDOG_HOST
it starts fine but logs connection errors.

Installation

The short way β€” one script detects the role (robot vs. brain), asks for the
other machine's address and sets everything up:

git clone https://github.com/rockywuest/pidog-embodiment.git && cd pidog-embodiment
sudo ./install.sh          # on the robot AND on the brain machine
# No robot? A dog made of log lines on :8888 β€” works with the MCP server too:
./install.sh --mock

(On the robot, install SunFounder's pidog SDK first β€” see the manual steps.)

The manual way, step by step:

# Clone the repo
git clone https://github.com/rockywuest/pidog-embodiment.git
cd pidog-embodiment

# === On the BODY (Pi 4 / Robot) ===
# First: install SunFounder's pidog + robot_hat SDKs with THEIR installer
# (https://github.com/sunfounder/pidog), they are not on PyPI-only.
cd body
# Raspberry Pi OS Bookworm enforces PEP 668 ("externally-managed-environment"):
pip3 install --break-system-packages -r requirements.txt
# Generates the systemd units for YOUR user and repo path, creates nox.env:
sudo ../scripts/install-body.sh
# Edit body/nox.env (set BRAIN_HOST β€” 127.0.0.1 if brain and body share one machine), then:
sudo systemctl start nox-body nox-bridge nox-voice

# === On the BRAIN (Pi 5 / Desktop) ===
# No pip install needed β€” the brain runs on the Python standard library only.
cd brain
# Generates the systemd unit for YOUR user/path and creates /etc/default/nox-brain:
sudo ../scripts/install-brain.sh
# Edit /etc/default/nox-brain: set PIDOG_HOST (127.0.0.1 if brain and body share
# one machine) and the LLM backend β€” OPENAI_API_KEY for OpenAI, or for a local
# Ollama: OPENAI_URL=http://127.0.0.1:11434/v1/chat/completions + LLM_MODEL=llama3.2
sudo systemctl start nox-brain

Note: Don't copy the services/*.service files verbatim β€” they contain a
reference user and paths. If you did and got
"failed because of unavailable resources or another system error", run the
install script above; it rewrites User= and all paths for your machine.

Ollama on a PC (CPU): local inference is slow β€” the brain defaults to a
120s LLM timeout for non-OpenAI endpoints (tune with LLM_TIMEOUT in
/etc/default/nox-brain). Small models like llama3.2 don't always return
the JSON the brain asks for; when that happens the dog speaks the raw reply
but performs no actions
. That's the graceful fallback, not a bug β€” a larger
model (e.g. llama3.1:8b, qwen2.5:7b) follows the JSON format much more
reliably.

Nox Mode vs. SunFounder Examples

The install script enables the body services, so Nox starts on every boot and
holds the robot's hardware (servos, touch, sound direction, ultrasonic). If you
then run SunFounder's own example scripts, their init fails with lines like
dual_touch init ... fail β€” both sides want exclusive access. Don't kill the
Python processes by hand (the SDK forks a helper child; killing the parent
leaves orphans). Switch modes via systemd instead:

# Run SunFounder examples (stops Nox until next boot):
sudo systemctl stop nox-body nox-bridge nox-voice

# Back to Nox mode:
sudo systemctl start nox-body nox-bridge nox-voice

# Make SunFounder mode survive reboots (Nox off at boot):
sudo systemctl disable nox-body nox-bridge nox-voice
# ...and to restore Nox autostart:
sudo systemctl enable --now nox-body nox-bridge nox-voice

Health Check

Not sure everything came up correctly? Run the doctor on the robot (and/or the
brain) β€” it auto-detects the role and checks services, the bridge API, the
battery, and the voice/vision models, with a fix hint for anything that's off:

./scripts/doctor.sh

If commands report ok but the dog doesn't physically move, run the servo
self-test (⚠️ it moves the dog β€” sit, stand, one-leg wiggle):

curl -s http://127.0.0.1:8888/selftest

It drives the servos directly, bypassing the SDK's queue/thread machinery, and
reports which user and PATH the daemon runs with, thread health, queue depth,
and β€” first of all β€” whether the robot_hat MCU answers on I2C ("i2c":
resolved address, devices actually on the bus, verdict and fix). The output
localizes the failing layer.

Symptom: every command answers ok, /status shows battery_v: 0.0, the
IMU init prints fail, the dog never moves β€” but SunFounder's examples work
under sudo.
That is the MCU not answering on I2C. robot_hat finds the MCU
(0x14/0x15/0x16, board-dependent) by running i2cdetect, which lives in
/usr/sbin; when that is not on the service's PATH the scan fails and the SDK
silently falls back to 0x14. Since the fix for issue #12 the units ship with
/usr/sbin on PATH, the daemon probes the MCU at startup and refuses motion
commands with an explicit error instead of a fake ok, and doctor.sh
checks the unit PATH and the bus. On an existing install:
git pull && sudo ./scripts/install-body.sh.

Where the Robot's Address Is Configured

No hostname is hard-coded in the code paths you run. Each side reads it from one
place:

What Where Setting
Body β†’ brain (callbacks, voice) body/nox.env BRAIN_HOST, BRAIN_CALLBACK_PORT
Brain voice service /etc/default/nox-brain (created by install-brain.sh) PIDOG_HOST
Examples and CLI environment or argument PIDOG_HOST, or ./scripts/deploy-body.sh user@host
# examples take the address without editing any file
PIDOG_HOST=192.168.1.42 python3 examples/basic_control.py
python3 examples/basic_control.py mydog.local

The I2C address of the robot_hat MCU is not configured at all: the SDK finds
it on the bus (0x14, 0x15 or 0x16) at every start. That is what the PATH fix in
#22 restored β€” see the troubleshooting note above.

Nox Mode vs. SunFounder Mode

The Nox services and SunFounder's own example scripts drive the same hardware
and cannot run at the same time (the examples fail with GPIO busy). Pick one:

# SunFounder examples now
sudo systemctl stop nox-body nox-bridge nox-voice

# boot into SunFounder mode by default
sudo systemctl disable nox-body nox-bridge nox-voice

# back to Nox mode, now and at every boot
sudo systemctl enable --now nox-body nox-bridge nox-voice

Inside Nox mode, the autonomous behavior engine (idle / patrol / play) starts
with the bridge. To keep the dog still unless you command it:

Scope How
Until the next restart curl -X POST http://127.0.0.1:8888/behavior/stop
Permanently add NOX_NO_AUTO=1 to body/nox.env, then sudo systemctl restart nox-bridge
On demand again curl -X POST http://127.0.0.1:8888/behavior/start -H 'Content-Type: application/json' -d '{"state": "idle"}'

Independent of the engine, the daemon protects the servos: after ~60 s without
commands the dog lies down, after ~120 s the servos switch off, and the next
command wakes it. The LEDs are the daemon's status light (purple awake, blue
resting, dim asleep) and keep breathing even with the engine disabled.

Known Upstream Quirks (worked around)

  • robot_hat 2.5.2a1: its get_battery_voltage() crashes with
    NameError: name '_adc_obj' is not defined. The daemon detects this and
    reads the battery ADC (channel A4) directly β€” /status stays correct.
  • robot_hat I2C layer swallows every bus error (retry, then return
    False) and locates the MCU via i2cdetect on PATH, falling back to 0x14
    when the scan yields nothing. A wrong address therefore looks like a healthy
    dog with a 0.0 V battery. The daemon probes the MCU itself, turns a dead bus
    into battery_v: "error" plus an I2C explanation, and fails motion commands
    loudly.
  • sound_effect init ... fail leaves Pidog without a .music object, and
    every sound path then raises AttributeError: 'Pidog' object has no attribute 'music'. Because bark, howling and pant call speak() inside their
    motion sequence, the movement was lost too (#24). The daemon now attaches an
    aplay-backed stand-in at startup, so those actions work with sound; the log
    says Audio: SDK sound engine missing, using aplay fallback.
  • SunFounder SDK do_action() silently ignores unknown actions and its
    action threads die permanently on their first exception. The daemon
    detects dead threads and reports them loudly instead of returning fake
    success.

Test It

# From the brain machine:
# Check robot status
curl http://your-robot.local:8888/status

# Make it sit (singular or an array of actions both work)
curl -X POST http://your-robot.local:8888/action \
  -H "Content-Type: application/json" \
  -d '{"action": "sit"}'

# Make it speak (answers at once, speaks in the background)
curl -X POST http://your-robot.local:8888/speak \
  -H "Content-Type: application/json" \
  -d '{"text": "Hallo! Ich bin online!"}'

# Silent? Wait for the speech and get the real outcome (piper or player error)
curl -X POST http://your-robot.local:8888/speak \
  -H "Content-Type: application/json" \
  -d '{"text": "Hallo!", "blocking": true}'

# Voice command (simulated) β€” needs the brain running (nox-brain on BRAIN_HOST).
# Answers {"ok": true, "brain": "received"}, or says why the brain was not reached.
curl -X POST http://your-robot.local:8888/voice/input \
  -H "Content-Type: application/json" \
  -d '{"text": "Setz dich hin und wedel mit dem Schwanz!"}'
# Without an LLM the brain understands keyword commands in German, English and
# French (sit/sitz/assis, lie/platz/couchΓ©, wag/wedel/remue la queue, …), several
# per sentence. With OPENAI_API_KEY or a local OPENAI_URL it understands anything.
# It answers in the language you spoke; NOX_LANG=fr (or de/en) in
# /etc/default/nox-brain fixes one β€” match it to the robot's Piper voice.

# Take a photo β€” ?format=jpeg returns the image itself (without it: JSON
# with the photo as base64 in "photo_b64", plus detected faces)
curl "http://your-robot.local:8888/photo?format=jpeg" -o snap.jpg

Voice for /speak: pip installs piper, but no voice. The daemon uses any
installed one β€” de_DE-thorsten-high if present, otherwise the first voice in
~/.local/share/piper-voices or ~/.piper_models (where SunFounder's examples
download theirs). A voice is a pair, <name>.onnx + <name>.onnx.json, from
rhasspy/piper-voices. To pick
one, set PIPER_MODEL=/full/path/to/voice.onnx in body/nox.env. If you ran
SunFounder's examples with sudo, their voices sit in /root/.piper_models,
out of the service's reach: sudo cp -r /root/.piper_models ~/ && sudo chown -R $USER: ~/.piper_models.
./scripts/doctor.sh shows which voice is used.

Shell quoting matters: wrap the JSON in single quotes and use
double quotes inside it, exactly as above. With the quotes swapped the
shell mangles the JSON before curl ever sends it β€” the bridge then answers
HTTP 400 with what it actually received, so you can see the mangling.

🎀 Talking to the Dog (voice input)

The nox-voice service (installed by default) listens on a USB microphone
for the wake word β€” "Nox" β€” and sends what you say to the brain, which answers
and acts. It stays off until a Vosk speech model is installed:

# 1. On the robot: a model for YOUR language decides what the dog understands.
#    Browse https://alphacephei.com/vosk/models β€” small models fit the Pi 4.
mkdir -p ~/vosk-models && cd ~/vosk-models
wget https://alphacephei.com/vosk/models/vosk-model-small-fr-0.22.zip && unzip vosk-model-small-*.zip

# 2. Point the service at the unzipped FOLDER, in body/nox.env:
#    VOSK_MODEL_PATH=/home/<you>/vosk-models/vosk-model-small-fr-0.22
sudo systemctl restart nox-voice
journalctl -u nox-voice -n 20 --no-pager    # expect the model to load, then "listening"

# 3. Say "Nox" ... then your command: "Assieds-toi !", "Sitz!", "sit down"

Without a model (or without a USB mic) nox-voice exits cleanly and voice
input stays off β€” ./scripts/doctor.sh tells you which of the two is missing.
Text commands via POST /voice/input work either way.

πŸ”Œ Claude as the Brain (MCP)

brain/nox_mcp_server.py exposes the robot as MCP tools β€” photo (the model
really sees the image), speak, move, emotions, sensors, emergency stop. Any MCP
client becomes the dog's brain; nox-brain is not needed for this.

# Claude Code (on your laptop β€” the robot just needs to be reachable):
claude mcp add pidog --env PIDOG_HOST=<robot-ip-or-name> --   python3 /path/to/pidog-embodiment/brain/nox_mcp_server.py
# then:  "take a photo, tell me what you see, and if a person is there, wag your tail"

Claude Desktop (claude_desktop_config.json) and other MCP clients:

{"mcpServers": {"pidog": {
  "command": "python3",
  "args": ["/path/to/pidog-embodiment/brain/nox_mcp_server.py"],
  "env": {"PIDOG_HOST": "<robot-ip-or-name>", "NOX_API_TOKEN": "<if auth is on>"}
}}}

Tools: dog_photo Β· dog_vision Β· dog_speak Β· dog_action Β· dog_look_at Β·
dog_expression Β· dog_rgb Β· dog_sensors Β· dog_status Β· dog_behavior
(incl. emergency_stop). Zero dependencies β€” plain stdlib, stdio transport.

πŸ“‘ API Reference

Bridge Endpoints (Body β€” Port 8888)

All 31 endpoints the bridge actually serves β€” see body/nox_brain_bridge.py.

Method Endpoint What it does
GET /status Battery, uptime, behavior state, security settings, last speech result
GET /sensors Ultrasonic distance, touch, battery, obstacle flags
GET /capabilities Valid actions, endpoints, feature flags
GET /selftest Servo/daemon self-test β€” physically moves the dog; reports dead threads, queue, audio
GET /photo Take a photo β€” JSON with photo_b64 + faces; ?format=jpeg returns the image itself
GET /look Photo + face detection + scene analysis
GET /vision Latest on-device SmolVLM scene description
GET /perception Current perception state (faces, objects, last photo)
GET /state Bridge-internal state snapshot
GET /faces Known faces in the recognition DB
GET /scan Ultrasonic distance scan
GET /scan/sweep Panoramic head-sweep scan
GET /memory/recent, /memory/stats Long-term memory reads (POST accepted too)
GET /voice/inbox Pending voice messages (polling reads clear it)
GET/POST /voice/echo_until TTS echo-suppression window
POST /speak TTS; "blocking": true waits and reports the real outcome
POST /action Execute action(s): {"action": "sit"} or {"actions": [...]}
POST /move Semantic movement: direction + distance_cm / angle_deg
POST /combo Actions + speech + RGB + head in one call
POST /expression Coordinated emotion: action + LEDs + head pose + sound
POST /rgb LED strip color and animation
POST /head Head pose (yaw/roll/pitch)
POST /look_at Head by direction (left/right/up/down/center) or angle/tilt
POST /face/register Register the face in front of the camera under a name
POST /voice/input Text command β†’ brain; reports whether the brain received it
POST /voice/respond Brain's reply β†’ spoken by the dog
POST /behavior/start Autonomous mode on (idle/patrol/play)
POST /behavior/stop Autonomous mode off (drains queued motion)
POST /emergency_stop Freeze NOW
POST /command Raw daemon passthrough (unvalidated by design)

Available Actions

Movement: forward, backward, turn_left, turn_right, stand, sit, lie, trot
Tricks:   wag_tail, bark, howling, pant, stretch, push_up, doze_off,
          hand_shake, high_five, scratch, body_twisting, lick_hand, feet_shake
Head:     nod, shake_head, tilting_head, think, recall
Posture:  attack_posture, sit_2_stand, waiting, alert, surprise

The exact set depends on your installed SunFounder SDK version: some are
ActionDict poses, others are preset functions β€” the daemon routes both
transparently. An unknown action returns an error listing everything your
SDK build actually supports
(also available via GET /capabilities).

Emotion β†’ RGB Mapping

{
  "happy":    {"r":0,   "g":255, "b":0,   "mode":"breath"},
  "sad":      {"r":0,   "g":0,   "b":128, "mode":"breath"},
  "curious":  {"r":0,   "g":255, "b":255, "mode":"breath"},
  "excited":  {"r":255, "g":255, "b":0,   "mode":"boom"},
  "alert":    {"r":255, "g":100, "b":0,   "mode":"boom"},
  "love":     {"r":255, "g":50,  "b":150, "mode":"breath"},
  "sleepy":   {"r":0,   "g":0,   "b":80,  "mode":"breath"}
}

πŸ”„ Multi-Body Support

The brain doesn't care what the body is β€” it talks HTTP. Switch bodies at runtime:

from brain.nox_body_client import BodyClient

# Connect to PiDog
dog = BodyClient("pidog.local", 8888)
dog.move("sit")
dog.speak("Ich bin ein Hund!")

# Switch to robot car
car = BodyClient("picar.local", 8888)
car.move("forward")
car.speak("Jetzt fahre ich!")

Adding a New Body

Implement the bridge API on your hardware:

# Minimum required endpoints:
POST /action    {"action": "forward|backward|left|right|stop"}
POST /speak     {"text": "..."}
GET  /status    β†’ {"battery_v": 7.4, "sensors": {...}}

See body/adapters/ for examples (PiDog, PiCar, custom).

🌐 Remote Access

Option 1: Tailscale (Recommended)

# On both brain and body:
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up

# Now use Tailscale IPs instead of .local addresses
export PIDOG_HOST="100.x.x.x"

Option 2: WireGuard

# See docs/remote-access.md for full WireGuard setup

Option 3: Telegram Bot

Control your robot from anywhere via Telegram:

export TELEGRAM_BOT_TOKEN="your-token"
# Required: without an allowlist the bot refuses every command. Send any message
# to the bot once β€” it replies with your user ID.
export TELEGRAM_ALLOWED_USERS="123456789"      # comma-separated for several
python3 brain/telegram_bot.py

Commands: /status, /photo, /speak <text>, /move <action>, /faces

πŸ‘οΈ Local Vision (SmolVLM-256M)

On-device scene understanding via llama.cpp β€” no cloud, no Python ML frameworks.

# Build llama.cpp on Pi 4 (one-time, ~20 min)
# Prerequisite: the build tools (skipping this gives "cmake: command not found")
sudo apt update && sudo apt install -y cmake build-essential
cd ~ && git clone --depth 1 https://github.com/ggml-org/llama.cpp.git
cd llama.cpp && cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_NEON=ON
cmake --build build --config Release -j2

# Download models (279 MB total; -c resumes instead of saving a second copy as .gguf.1)
mkdir -p ~/models/smolvlm && cd ~/models/smolvlm
wget -c https://huggingface.co/ggml-org/SmolVLM-256M-Instruct-GGUF/resolve/main/SmolVLM-256M-Instruct-Q8_0.gguf
wget -c https://huggingface.co/ggml-org/SmolVLM-256M-Instruct-GGUF/resolve/main/mmproj-SmolVLM-256M-Instruct-Q8_0.gguf

# Start the vision service β€” it is not part of the default install
cd ~/pidog-embodiment && sudo ./scripts/install-body.sh nox-vision
journalctl -u nox-vision -n 20 --no-pager   # expect "[vision] Setup OK"

# About a minute later: check what PiDog sees
curl -s http://your-robot.local:8888/vision | python3 -m json.tool

"vision not running (no result file)" means the service never wrote a result β€”
it is not installed or not running. Any other error in the answer is the
service's own reason (missing model, camera, llama.cpp).

Performance on Pi 4 (2GB RAM):

Metric Value
Inference time ~27s (warm) / ~37s (cold)
Generation speed ~3.2 tokens/sec
RAM usage ~400MB peak
Model size 279 MB (167 + 112 MB)

πŸ‘€ Face Recognition

Uses SCRFD (detection) + ArcFace (recognition) via ONNX Runtime. Runs on the body (Pi 4).

# Download ONNX models (one-time)
cd models
./download_models.sh

# Register a face via API
curl -X POST http://your-robot.local:8888/face/register \
  -H "Content-Type: application/json" \
  -d '{"name": "Rocky"}'

# Identify faces in current view
curl -X POST http://your-robot.local:8888/face/identify

# Performance (Pi 4):
# Detection: ~400ms | Embedding: ~188ms | Full: ~567ms

πŸ€– Autonomous Behaviors

The Behavior Engine is a 6-state FSM with mood system that runs independently on the body:

States

  • Idle β†’ Random head movements, occasional tail wag, energy recovery
  • Patrol β†’ Autonomous navigation with ultrasonic + vision obstacle avoidance
  • Investigate β†’ Approach detected person/sound, face tracking
  • Alert β†’ Threat response (bark, red LEDs, report to brain)
  • Play β†’ Interactive play when touched (tail wag, happy LEDs, tricks)
  • Rest β†’ Low-power state, minimal movement, PWM auto-disable

Built-in Reflexes (work without brain)

  • Touch β†’ Pat on head triggers tail wag + happy LEDs
  • Sound β†’ Head turns toward sound source
  • Battery β†’ Warning and rest below 6.8 V; below 6.2 V it lies down and stays down until charged
  • Vision β†’ Patrol uses SmolVLM to detect people and obstacles
  • Face tracking β†’ Head follows detected faces
# Start patrol mode
curl -X POST http://your-robot.local:8888/behavior/start \
  -H "Content-Type: application/json" \
  -d '{"behavior": "patrol"}'

# Stop all behaviors (servos auto-disable after 120s idle)
# Also drops motion frames already queued, so the dog stops now instead of
# finishing the queue; the reply reports how many were dropped (issue #25).
curl -X POST http://your-robot.local:8888/behavior/stop
# β†’ {"ok": true, "stopped": true, "motion_frames_dropped": 227}

πŸ›‘οΈ Security β€” read this before exposing the robot

The bridge has no authentication. Anyone who can reach port 8888 can walk
the robot, take photos and play sounds. Treat it as a LAN-only service.

Status
Bridge authentication βœ… optional: set NOX_API_TOKEN in body/nox.env and every request from another machine needs Authorization: Bearer <token>. Off by default so existing setups keep working. The robot's own services (localhost) are exempt; requests through a tunnel/proxy on the robot are not. /status β†’ security.auth shows on/off. Treat the token like a password: the API allows any origin (CORS *), so the token is the only thing standing between a browser on your LAN and the robot.
Rate limiting βœ… 600 requests/minute per remote address (NOX_RATE_LIMIT, 0 = off), answered with 429 + Retry-After
Input validation βœ… unknown actions and malformed JSON are rejected loudly; /rgb, /head, /speak and /face/register check their values and answer 400 with the reason
Secrets in code βœ… none β€” everything comes from the environment, enforced by tests/test_no_private_data.py
Telegram bot βœ… refuses every command unless TELEGRAM_ALLOWED_USERS lists your user ID

Recommended setup:

# 1. Do not forward port 8888 from your router.
# 2. Restrict the bridge to the LAN and your VPN (example for ufw):
sudo ufw allow from 192.168.0.0/16 to any port 8888 proto tcp
sudo ufw allow from 100.64.0.0/10 to any port 8888 proto tcp   # Tailscale CGNAT
# 3. Reach the robot from outside via VPN only β€” see docs/remote-access.md.
# 4. Set a token on both sides (body/nox.env and /etc/default/nox-brain):
#    NOX_API_TOKEN=$(python3 -c "import secrets; print(secrets.token_urlsafe(24))")

Ports in use: 8888 inbound on the body (bridge), 8889 inbound on the
brain
(voice callbacks), and 9999 on the body's loopback only (daemon).

πŸ“ Project Structure

pidog-embodiment/
β”œβ”€β”€ brain/                         # Runs on Pi 5 / Desktop
β”‚   β”œβ”€β”€ nox_body_client.py         # Python client for bridge API (37 functions)
β”‚   β”œβ”€β”€ nox_voice_brain.py         # LLM-powered voice processing
β”‚   β”œβ”€β”€ nox_voice_relay.py         # Voice relay for remote STT
β”‚   β”œβ”€β”€ nox_mcp_server.py          # The robot as MCP tools (Claude Code/Desktop, OpenClaw)
β”‚   β”œβ”€β”€ telegram_bot.py            # Telegram remote control
β”‚   β”œβ”€β”€ requirements.txt
β”‚   └── services/
β”‚       └── nox-brain.service
β”œβ”€β”€ body/                          # Runs on Pi 4 / Robot
β”‚   β”œβ”€β”€ nox_daemon.py              # Low-level hardware daemon (servos, sensors, camera)
β”‚   β”œβ”€β”€ nox_brain_bridge.py        # HTTP REST API server (20+ endpoints)
β”‚   β”œβ”€β”€ nox_behavior_engine.py     # 6-state FSM + mood system + obstacle avoidance
β”‚   β”œβ”€β”€ nox_vision.py              # Local vision engine (SmolVLM-256M via llama.cpp)
β”‚   β”œβ”€β”€ nox_face_recognition.py    # SCRFD detection + ArcFace recognition
β”‚   β”œβ”€β”€ pidog_memory.py            # Drift-style memory with co-occurrence + decay
β”‚   β”œβ”€β”€ nox_voice_loop_v2.py       # Wake word + Vosk STT (what nox-voice runs)
β”‚   β”œβ”€β”€ nox_voice_loop_v3.py       # Experimental: faster-whisper + VAD variant
β”‚   β”œβ”€β”€ nox_control.py             # Direct servo control utilities
β”‚   β”œβ”€β”€ nox_i2c_diag.py            # I2C/MCU reachability diagnostics (issue #12)
β”‚   β”œβ”€β”€ nox_motion.py              # Draining the SDK's motion queue (issue #25)
β”‚   β”œβ”€β”€ nox_audio.py               # aplay fallback when the SDK has no sound (issue #24)
β”‚   β”œβ”€β”€ nox_security.py            # Bridge token auth, rate limit, input checks (issue #30)
β”‚   β”œβ”€β”€ adapters/                  # Hardware-specific adapters
β”‚   β”‚   β”œβ”€β”€ pidog.py               # SunFounder PiDog
β”‚   β”‚   └── picar.py               # Robot car (template)
β”‚   β”œβ”€β”€ requirements.txt
β”‚   └── services/
β”‚       β”œβ”€β”€ nox-body.service       # Hardware daemon (TCP 9999)
β”‚       β”œβ”€β”€ nox-bridge.service     # REST API (HTTP 8888)
β”‚       β”œβ”€β”€ nox-vision.service     # Vision engine (SmolVLM)
β”‚       └── nox-voice.service      # Wake word + STT
β”œβ”€β”€ models/                        # ONNX + GGUF models (gitignored)
β”‚   └── download_models.sh         # One-click model download
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ install-body.sh            # Install the body services for this machine
β”‚   β”œβ”€β”€ install-brain.sh           # Install the brain service
β”‚   β”œβ”€β”€ deploy-body.sh             # Copy body code to a robot over SSH
β”‚   β”œβ”€β”€ doctor.sh                  # Post-install health check
β”‚   └── pidog.sh                   # CLI control script
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ architecture.md
β”‚   β”œβ”€β”€ remote-access.md
β”‚   └── media-guide.md
β”œβ”€β”€ tests/                         # pytest β€” runs without hardware
β”œβ”€β”€ examples/
β”‚   β”œβ”€β”€ basic_control.py
β”‚   β”œβ”€β”€ face_registration.py
β”‚   └── multi_body.py
└── README.md

🀝 Contributing

This project is open source! We'd love contributions for:

  • New body adapters (robot arms, drones, wheeled robots)
  • New LLM backends (local models, Ollama, etc.)
  • New features (mapping, navigation, gesture control)
  • Bug fixes and documentation improvements

πŸ“œ License

MIT License β€” use it, modify it, build cool robots with it.

β˜• Support

If this project helped you or made you smile, consider buying us a coffee:

Ko-fi

πŸ’‘ Inspiration

"Every AI deserves a body to explore the world with."

Built by Nox ⚑ (an AI running on Clawdbot) and Rocky.


⭐ Star this repo if you want your AI to have legs!

Yorumlar (0)

Sonuc bulunamadi