At a glance
Measured on lab recordings and live robot runs from September to October 2026. Details and conditions are in Results.
What's new
Latest changes in this repository. Every masking step has a switch in .env, and the 1 Oct version is kept on branch v1-2026-10-01.
- Held rule is the default. An object touching a person is removed only if it moves or is held in the hand. It ran live on 5 and 6 Oct: a bottle in the hand was removed, every suitcase, laptop and standing bottle beside a person was kept.
- Masking v9, tested live on
map803_1006_run2. In live runs people stood against walls and glass, and the 1 Oct box fill removed walls and bags with them. The box fill is replaced by a narrow leg fill, every step is limited to the area around a person, static objects are protected, carving uses the camera's real depth, and a thinner fused map is written. Result: 85 % of removed pixels are people, alignment 0.104 m. Details - Two fixes. The box fill no longer swallows a suitcase standing beside a person, and a second recording on a running server no longer crashes.
- Person points left in the 3D map: 36,244 → 3,546. Five changes against people leaning on stools and boxes: mask dilation, filling the person's box at their depth (and below it), a 3-frame bridge, a depth-edge ring, and multi-view 3D carving. Also a fix for the overlap frame shared by two submaps. Kept on branch
v1-2026-10-01. SEMANTIC=1inrun_t1_server.shbuilds the semantic instance map and the deploy files in the same run, so alignment and the navigation bundle work with masking on.- SAM 3.1 compared fairly with YOLO. An earlier SAM 3 result came from a class-id bug; after the fix the two detectors find people in almost the same frames (117 vs 116). A stricter "carried" rule was added as an option (off by default).
- Geo gate. Motion-only blobs must touch a detected movable object. Person-free frames with >1 % removed: 59 → 1 of 296.
- Live mask view. The 3D viewer shows removed points in red, and every frame is saved with its mask to
mask_viz/.
How the whole system works
One command, run_bridge_oneshot.sh, runs the whole mapping session. It stops twice to ask the operator before anything moves.
Coverage route
- 2D map from the Kachaka Appexported as PNG + YAML, map ID checked
- Coverage pathboustrophedon or spiral, robot-radius and wall-clearance checks
- Goals + previewthinned waypoints, yaw limits, ordered preview image
goals_RUN.csv, preview PNGRobot + camera
- Kachaka
move_to_posegoal by goal, onboard obstacle avoidance, 60 s timeout per goal - Pose2D loggerrobot pose in the 2D map frame
- RealSense D435 → ROS 2 gatewayRGB + depth + intrinsics streamed to the SLAM server
pose2d_RUN.csv, RGB-D stream3D SLAM + masking
- SLAM serversubmaps of 16 frames, loop closure, live 3D view
- Dynamic maskingper chunk, between depth inference and point creation
- 3D carvingremoves points that other views see on a person
static_only_pcd.ply, mask_viz/Map for navigation
- Sim(2) alignmentSLAM trajectory ↔ Pose2D log, camera lever arm corrected
- Semantic instancesobjects with positions in the robot's map frame (
SEMANTIC=1) - Navigation bundleoccupancy + objects for the robotic agent
The seven stages of a run
| Stage | What happens | Operator |
|---|---|---|
| 1 / 7 | Preflight (containers, camera, arm, disk, stale processes) and show the active Kachaka map | Confirm the map |
| 2 / 7 | Export the 2D map to artifacts/runs/kachaka_2d_VENUE.{png,yaml} | |
| 3 / 7 | Plan the coverage path, thin it to goals, render the preview | Check the preview |
| 4 / 7 | Start the Pose2D logger | |
| 5 / 7 | Start 3D SLAM with masking (1–3 min to load models) | |
| 6 / 7 | Drive the goals one by one; a goal over the timeout is cancelled and skipped | Stand at the e-stop |
| 6.5 | Recording finishes; the map is saved | Never kill the server |
| 7 / 7 | Print (or run with --align) the alignment commands |
Dynamic-object masking
People are always removed. Other movable objects (laptop, bag, bottle…) are removed only while they move or are held, so a laptop left on a desk or a suitcase beside someone stays in the map. Masked pixels never become 3D points.
Semantic: YOLOv9e-seg
Instance masks per frame. Person confidence ≥ 0.15 (others ≥ 0.25), set low on purpose for blurry robot footage.
A person is always removed. Another movable object is removed only if it is moving (its pixels move against their surroundings and its 3D centre moves ≥ 10 cm) or held: it touches a person, lies mostly inside their outline, at their depth (±25 cm), and is small next to them.
Motion: FlowSeek optical flow
Measured flow is compared with the flow that the camera's own motion predicts. Pixels whose residual exceeds an adaptive threshold (median + 3·MAD), checked against both neighbouring frames, are moving.
Geo gate: a motion blob is kept only next to a detected person (30 px).
Bridge
Fills up to 3 frames without a detection when the same chunk has a detection before and after the gap, so a few blurry frames do not leave a "ghost" in the map.
It only fills pixels at the depth of the person it was carried from (±0.3 m).
DYNAMIC_PERSON_DILATE_PX=5DYNAMIC_LEG_FILL=1DYNAMIC_BRIDGE_MAX_GAP=3DYNAMIC_EDGE_RING_PX=8carved_pcd.ply.DYNAMIC_CARVE=1
lab_20260925, run move1). Top-down view of the 3D map before (left) and after (right) removal. The person-shaped smears in the open floor area disappear; walls and furniture are kept.Watching it live
While the robot maps, a web view shows the 3D map with removed points in red next to the camera panel, colour-coded by channel (semantic, motion only, bridge). Every frame is also saved to mask_viz/ with a per-frame removed.csv.
Results
All runs are in our lab with the Kachaka robot. "Offline" means a recording replayed through exactly the same server code path as a live run.
| Run | Type | What was measured | Result |
|---|---|---|---|
mapping_20260922_01 | Offline, 389 frames | Mean pixels removed per frame / motion-only share | 2.2 % / 0.1 % |
mapping_20260922_01 | Offline | Person-free frames with >1 % removed, before → after geo gate | 59 → 1 of 296 |
map803_0930_run5 | Live + semantic | Person points left in the 3D map (1 Oct version; v9: about 6.2k) | 36,244 → 3,546 |
map803_0930_run5 | Live + semantic | Alignment RMSE (lever arm corrected), max error | 0.077 m, 0.151 m |
map803_0930_run5 | Nav bundle | Objects with a goal / in mapped space | 57 / 68, 25 |
map803_1006_run2 | Live, 251 frames (v9) | Share of removed pixels that are people; basket and suitcase | 85 %, kept |
map803_1006_run2 | Live, 251 frames (v9) | Alignment RMSE | 0.104 m pass |
onsite_remap_run3 | Live, 407 frames | Alignment RMSE (reference baseline) | 0.104 m pass |
ec129_0911_bridge | Live, 910 frames | Alignment RMSE | 0.108 m pass |
lab_20260925 run3, move1 | Live, people walking | Walking, sitting and carrying people removed | qualitative |
Detector comparison: YOLOv9e-seg vs SAM 3.1
Same recording, 413 chunk-frames. SAM 3.1 runs only offline (≈0.7 s/frame on our Jetson), so YOLO stays the live detector.
| Detector | Frames with a person | Final mask area | Removed on frames neither detector sees |
|---|---|---|---|
| YOLOv9e-seg + geo gate live | 116 | 2.37 % | 0.002 % |
| SAM 3.1 + geo gate | 117 | 2.09 % | 0.008 % |
The two detectors agree on 105 frames; their union covers 128. On the 6 Oct runs YOLO already finds the person in 98–99 % of the reference pixels, and a YOLO + SAM 3.1 ensemble gains little: the person points that remain come from depth spill around people, not from missed detections.
Getting started
Planning runs on any computer with Python 3.10+. The live pipeline needs the lab machine (Docker, GPU, the SLAM and gateway images, model weights, the Kachaka SDK).
git clone https://github.com/Gauravmeena1/dynamic_SLAM.git
cd dynamic_SLAM
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m unittest discover -s . -p 'test_*.py' -v
# plan a route on any occupancy map (no robot, nothing moves)
python make_bridge_traj.py --map-yaml map.yaml --start 0,0,0 \
--pattern boustrophedon --spacing 0.6 --min-clearance 0.40 \
--out artifacts/runs/traj_demo.csv
python drive_waypoints.py artifacts/runs/traj_demo.csv \
--map-yaml map.yaml --save-goals artifacts/runs/goals_demo.csv
python preview_traj.py --map-yaml map.yaml --traj artifacts/runs/goals_demo.csvWithout --go, drive_waypoints.py only prints the plan.
cp .env.example .env && nano .env # robot address, camera serial, paths, container names # model weights are not in git: yolov9e-seg.pt and thirdparty/ (FlowSeek) cp <method-folder>/yolov9e-seg.pt dynamic_masking/ cp -r <method-folder>/thirdparty dynamic_masking/ set -a; . ./.env; set +a slam_integration/install.sh /path/to/ma-long-server "$KACHAKA_TOOLS_DIR" slam_tools/make_containers.sh # only creates missing containers # after every reboot docker start "$KACHAKA_AI_CONTAINER" "$KACHAKA_SLAM_CONTAINER" $KACHAKA_PYTHON slam_tools/kachaka_ctl.py status slam_tools/cam_check.sh # camera must report USB 3.x ./preflight.sh --name RUN
Every .env key is described in the masking guide. Ask Gaurav for the weights folder on the lab machine.
# window 1: SLAM server with masking (wait for "models resident") tmux new -s t1 slam_tools/run_t1_server.sh --mode rgb+depth+intr --backend ma # window 2: replay a recording through the same code path as live python3 slam_tools/replay_ab.py mapping_20260922_01 offline_test1 # top-down before | after | removed docker exec "$KACHAKA_AI_CONTAINER" python \ "$KACHAKA_CODE_DIR_CT/dynamic_SLAM/slam_tools/bev_compare.py" /fungi/outputs_malong/offline_test1
Expected on mapping_20260922_01: about 2.2 % of pixels removed per frame, about 0.1 % from motion alone (measured with the September method). For the September behaviour, put DYNAMIC_PERSON_DILATE_PX=0 DYNAMIC_LEG_FILL=0 DYNAMIC_BRIDGE_MAX_GAP=1 DYNAMIC_EDGE_RING_PX=0 DYNAMIC_CARVE=0 DYNAMIC_CARRIED_RULE=touch in .env; for the 1 Oct version, use branch v1-2026-10-01.
# 1. plan only, then open artifacts/previews/preview_goals_RUN.png ./run_bridge_oneshot.sh --name RUN --venue VENUE --goal-timeout 60 --plan-only # 2. full run in tmux: answer y twice (map correct? / route OK, robot moves?) tmux new -s mapping ./run_bridge_oneshot.sh --name RUN --venue VENUE --reuse-map --goal-timeout 60 # semantic instances + deploy files (needed for alignment and the nav bundle): # add SEMANTIC=1 to .env before the run; every slam_tools script loads .env echo SEMANTIC=1 >> .env # 3. watch live python3 slam_tools/live_view.py "$KACHAKA_TOOLS_DIR/outputs_malong/RUN" --mask --port 8092
The DYNAMIC_* switches go in .env the same way. Use a new --name every run and a new --venue whenever the App map changes. Start about 1 m off the dock; planning from the dock fails.
--plan-only first, check the preview, keep the path clear, and keep a trained person within reach of the emergency stop. Don't use -y on a first run. Never kill the SLAM server while it records or saves, because the map in memory is lost.Repository layout
Machine-specific values (robot address, camera serial, paths) live only in a local .env and are never committed.
| Path | Purpose |
|---|---|
run_bridge_oneshot.sh | The seven-stage field workflow, with confirmation before motion |
make_bridge_traj.py, drive_waypoints.py, preview_traj.py | Coverage planning, goal thinning and validation, previews |
preflight.sh, arm_cam_tune.sh | Read-only site checks; camera view and arm stow-pose tuning |
dynamic_masking/ | The masking method: chunk_fusion_masker.py (fusion, object policy, geo gate), dynamic_object_mask.py (YOLO), dynamic_fusion.py + flowseek_flow.py (motion), motion_compensation.py |
slam_integration/ | fusion_solver.py hooks the masker into the SLAM server, adds the live view and 3D carving; small patches for the server and camera gateway; idempotent install.sh |
slam_tools/ | Server and gateway launchers, live view, replay, robot control (kachaka_ctl.py), camera check, cleanup, bird's-eye comparison |
docs/ | Guides, validation reports and this page |
Hardware and software
- Robot: Kachaka mobile base (driven through the Kachaka API,
move_to_pose), with a robot arm that holds the camera in a fixed stow pose during mapping. - Camera: Intel RealSense D435 RGB-D on USB 3.
- Compute: NVIDIA Jetson (aarch64) running two Docker containers: a camera-to-ROS 2 gateway and a GPU container with the SLAM server and the masker.
- Models: YOLOv9e-seg (semantic), FlowSeek (optical flow), the lab's ma-long SLAM backend; SAM 3.1 for offline evaluation.
Documentation
All documents are in the repository.
Limitations and next steps
Known limitations
- With the geo gate on, an object the detector cannot label is no longer removed by motion alone (
DYNAMIC_GEO_GATE=nonerestores it). - A few thousand person points remain per run, from depth spill at a person's edges and legs behind furniture.
- In blurred frames while the robot turns, the bridge can still paste a person shape onto a static object.
- People far away or behind frosted glass are placed at the wrong depth by the model.
- The held-rule thresholds come from only a few recordings.
- Tinted glass lets depth see through it, which gives blotchy areas; this isn't a masking error.
- A 0.40 m wall clearance removes much of the free space, so always check the preview.
Being evaluated
- Giving YOLO its expected BGR colour order: on run2 it cut frames with a missed person from 11 to 3 (offline).
- Stopping the bridge in blurred turning frames.
- Using the robot's LiDAR pose as a prior to reduce SLAM drift.