Kachaka robot · RGB-D SLAM · dynamic-object removal

Automatic 3D mapping with people removed from the map

Kachaka 自動建圖 · 3D SLAM · 即時動態物件移除

The Kachaka robot drives a planned coverage route while an RGB-D camera records. A 3D SLAM server builds the map and, at the same time, removes people and moving objects so they never become map points. The 3D map is then aligned to the robot's 2D map and packaged for the robotic agent's navigation.

Four pairs of camera frames: raw frame on the left, the same frame with the person and a carried bottle covered by a red removal mask on the right
Live removal during mapping. Each pair shows the raw RGB frame and the pixels the masker removed (red = person/semantic channel). Walking people, close-ups, motion blur and a carried bottle are all removed before the frame reaches the map.

At a glance

Measured on lab recordings and live robot runs from September to October 2026. Details and conditions are in Results.

0.077 m
2D/3D alignment RMSE on the map803 run (pass threshold 0.30 m)
59 → 1
person-free frames with >1 % wrongly removed (of 296) after the geo gate
85 %
of removed pixels are people on the 6 Oct live run; basket and suitcase kept
57 / 68
semantic objects that get a reachable navigation goal for the robotic agent

What's new

Latest changes in this repository. Every masking step has a switch in .env, and the 1 Oct version is kept on branch v1-2026-10-01.

  1. Held rule is the default. An object touching a person is removed only if it moves or is held in the hand. It ran live on 5 and 6 Oct: a bottle in the hand was removed, every suitcase, laptop and standing bottle beside a person was kept.
  2. Masking v9, tested live on map803_1006_run2. In live runs people stood against walls and glass, and the 1 Oct box fill removed walls and bags with them. The box fill is replaced by a narrow leg fill, every step is limited to the area around a person, static objects are protected, carving uses the camera's real depth, and a thinner fused map is written. Result: 85 % of removed pixels are people, alignment 0.104 m. Details
  3. Two fixes. The box fill no longer swallows a suitcase standing beside a person, and a second recording on a running server no longer crashes.
  4. Person points left in the 3D map: 36,244 → 3,546. Five changes against people leaning on stools and boxes: mask dilation, filling the person's box at their depth (and below it), a 3-frame bridge, a depth-edge ring, and multi-view 3D carving. Also a fix for the overlap frame shared by two submaps. Kept on branch v1-2026-10-01.
  5. SEMANTIC=1 in run_t1_server.sh builds the semantic instance map and the deploy files in the same run, so alignment and the navigation bundle work with masking on.
  6. SAM 3.1 compared fairly with YOLO. An earlier SAM 3 result came from a class-id bug; after the fix the two detectors find people in almost the same frames (117 vs 116). A stricter "carried" rule was added as an option (off by default).
  7. Geo gate. Motion-only blobs must touch a detected movable object. Person-free frames with >1 % removed: 59 → 1 of 296.
  8. Live mask view. The 3D viewer shows removed points in red, and every frame is saved with its mask to mask_viz/.

How the whole system works

One command, run_bridge_oneshot.sh, runs the whole mapping session. It stops twice to ask the operator before anything moves.

1 · Plan

Coverage route

  1. 2D map from the Kachaka Appexported as PNG + YAML, map ID checked
  2. Coverage pathboustrophedon or spiral, robot-radius and wall-clearance checks
  3. Goals + previewthinned waypoints, yaw limits, ordered preview image
→ goals_RUN.csv, preview PNG
2 · Drive & record

Robot + camera

  1. Kachaka move_to_posegoal by goal, onboard obstacle avoidance, 60 s timeout per goal
  2. Pose2D loggerrobot pose in the 2D map frame
  3. RealSense D435 → ROS 2 gatewayRGB + depth + intrinsics streamed to the SLAM server
→ pose2d_RUN.csv, RGB-D stream
3 · Build

3D SLAM + masking

  1. SLAM serversubmaps of 16 frames, loop closure, live 3D view
  2. Dynamic maskingper chunk, between depth inference and point creation
  3. 3D carvingremoves points that other views see on a person
→ static_only_pcd.ply, mask_viz/
4 · Align & export

Map for navigation

  1. Sim(2) alignmentSLAM trajectory ↔ Pose2D log, camera lever arm corrected
  2. Semantic instancesobjects with positions in the robot's map frame (SEMANTIC=1)
  3. Navigation bundleoccupancy + objects for the robotic agent
→ alignment report, nav bundle

The seven stages of a run

StageWhat happensOperator
1 / 7Preflight (containers, camera, arm, disk, stale processes) and show the active Kachaka mapConfirm the map
2 / 7Export the 2D map to artifacts/runs/kachaka_2d_VENUE.{png,yaml}
3 / 7Plan the coverage path, thin it to goals, render the previewCheck the preview
4 / 7Start the Pose2D logger
5 / 7Start 3D SLAM with masking (1–3 min to load models)
6 / 7Drive the goals one by one; a goal over the timeout is cancelled and skippedStand at the e-stop
6.5Recording finishes; the map is savedNever kill the server
7 / 7Print (or run with --align) the alignment commands
Occupancy map with a planned serpentine route of 17 waypoints from a blue start dot to a red end dot
Stage 3 preview. 17 goals, 10.2 m, blue = start, red = end, arrows show direction. The operator checks this before the robot moves.
2D map with the robot's logged path in green and the aligned SLAM camera path in blue, plus labelled semantic objects
Stage 7 alignment. Robot path from the Pose2D log and SLAM camera path after Sim(2) alignment, with semantic objects placed in the robot's map. The visible gap is the camera's 0.146 m lever arm; with it corrected the RMSE is 0.077 m.

Dynamic-object masking

People are always removed. Other movable objects (laptop, bag, bottle…) are removed only while they move or are held, so a laptop left on a desk or a suitcase beside someone stays in the map. Masked pixels never become 3D points.

Semantic: YOLOv9e-seg

Instance masks per frame. Person confidence ≥ 0.15 (others ≥ 0.25), set low on purpose for blurry robot footage.

A person is always removed. Another movable object is removed only if it is moving (its pixels move against their surroundings and its 3D centre moves ≥ 10 cm) or held: it touches a person, lies mostly inside their outline, at their depth (±25 cm), and is small next to them.

Motion: FlowSeek optical flow

Measured flow is compared with the flow that the camera's own motion predicts. Pixels whose residual exceeds an adaptive threshold (median + 3·MAD), checked against both neighbouring frames, are moving.

Geo gate: a motion blob is kept only next to a detected person (30 px).

Bridge

Fills up to 3 frames without a detection when the same chunk has a detection before and after the gap, so a few blurry frames do not leave a "ghost" in the map.

It only fills pixels at the depth of the person it was carried from (±0.3 m).

Person cleanup (v9, Oct 6). YOLO's mask stops at an occluder, depth edges leave "flying" pixels, and other frames' predicted depth can put a point where a person stands. Five steps, all on by default; each one only acts close to a person, so walls, bags and furniture next to people stay:
DilateGrow every person mask by 5 px.
DYNAMIC_PERSON_DILATE_PX=5
Leg fillBelow a person's outline, when something stands in front, add pixels at lower-body depth (never the occluder, never a detected object).
DYNAMIC_LEG_FILL=1
Bridge 3 framesFill short detector gaps at the person's depth.
DYNAMIC_BRIDGE_MAX_GAP=3
Edge ringDrop depth-jump pixels within 8 px of a person.
DYNAMIC_EDGE_RING_PX=8
3D carvingUsing the camera's real depth, drop map points that frames only ever see on removed pixels; saved to carved_pcd.ply.
DYNAMIC_CARVE=1
Two frames without people. Left: yellow motion-only masks on a door and shelf tops. Right: the same frames with nothing removed.
Geo gate removes false positives. Left: before, the flow channel fired on a door edge and shelf tops while the camera turned (yellow = motion only). Right: after, nothing is removed. Person-free frames with >1 % removed fell from 59/296 to 1/296; the people removal was unchanged.
Point-cloud renders: before, magenta person points visible; after, almost none
3D carving removes leftover person points. Renders of the 3D map from two camera poses; magenta = points that lie on a person, blue = person outline. With the 1 Oct version, person points across the run fell from 36,244 to 3,546; v9 keeps about 6.2k, because it no longer removes the walls and objects around people.
Top-down views of a lab 3D map before and after removal; the after view lacks the scattered person-shaped blobs in the middle of the room
Live run with people walking through (map lab_20260925, run move1). Top-down view of the 3D map before (left) and after (right) removal. The person-shaped smears in the open floor area disappear; walls and furniture are kept.

Watching it live

While the robot maps, a web view shows the 3D map with removed points in red next to the camera panel, colour-coded by channel (semantic, motion only, bridge). Every frame is also saved to mask_viz/ with a per-frame removed.csv.

Results

All runs are in our lab with the Kachaka robot. "Offline" means a recording replayed through exactly the same server code path as a live run.

RunTypeWhat was measuredResult
mapping_20260922_01Offline, 389 framesMean pixels removed per frame / motion-only share2.2 % / 0.1 %
mapping_20260922_01OfflinePerson-free frames with >1 % removed, before → after geo gate59 → 1 of 296
map803_0930_run5Live + semanticPerson points left in the 3D map (1 Oct version; v9: about 6.2k)36,244 → 3,546
map803_0930_run5Live + semanticAlignment RMSE (lever arm corrected), max error0.077 m, 0.151 m
map803_0930_run5Nav bundleObjects with a goal / in mapped space57 / 68, 25
map803_1006_run2Live, 251 frames (v9)Share of removed pixels that are people; basket and suitcase85 %, kept
map803_1006_run2Live, 251 frames (v9)Alignment RMSE0.104 m pass
onsite_remap_run3Live, 407 framesAlignment RMSE (reference baseline)0.104 m pass
ec129_0911_bridgeLive, 910 framesAlignment RMSE0.108 m pass
lab_20260925 run3, move1Live, people walkingWalking, sitting and carrying people removedqualitative

Detector comparison: YOLOv9e-seg vs SAM 3.1

Same recording, 413 chunk-frames. SAM 3.1 runs only offline (≈0.7 s/frame on our Jetson), so YOLO stays the live detector.

DetectorFrames with a personFinal mask areaRemoved on frames neither detector sees
YOLOv9e-seg + geo gate live1162.37 %0.002 %
SAM 3.1 + geo gate1172.09 %0.008 %

The two detectors agree on 105 frames; their union covers 128. On the 6 Oct runs YOLO already finds the person in 98–99 % of the reference pixels, and a YOLO + SAM 3.1 ensemble gains little: the person points that remain come from depth spill around people, not from missed detections.

Two views of a coloured 3D point-cloud map of a lab room with the ceiling removed
A finished 3D map of a lab room (ceiling removed), bird's-eye and oblique views.

Getting started

Planning runs on any computer with Python 3.10+. The live pipeline needs the lab machine (Docker, GPU, the SLAM and gateway images, model weights, the Kachaka SDK).

git clone https://github.com/Gauravmeena1/dynamic_SLAM.git
cd dynamic_SLAM
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python -m unittest discover -s . -p 'test_*.py' -v

# plan a route on any occupancy map (no robot, nothing moves)
python make_bridge_traj.py --map-yaml map.yaml --start 0,0,0 \
  --pattern boustrophedon --spacing 0.6 --min-clearance 0.40 \
  --out artifacts/runs/traj_demo.csv
python drive_waypoints.py artifacts/runs/traj_demo.csv \
  --map-yaml map.yaml --save-goals artifacts/runs/goals_demo.csv
python preview_traj.py --map-yaml map.yaml --traj artifacts/runs/goals_demo.csv

Without --go, drive_waypoints.py only prints the plan.

⚠ SafetyA full run moves the robot. Always run --plan-only first, check the preview, keep the path clear, and keep a trained person within reach of the emergency stop. Don't use -y on a first run. Never kill the SLAM server while it records or saves, because the map in memory is lost.

Repository layout

Machine-specific values (robot address, camera serial, paths) live only in a local .env and are never committed.

PathPurpose
run_bridge_oneshot.shThe seven-stage field workflow, with confirmation before motion
make_bridge_traj.py, drive_waypoints.py, preview_traj.pyCoverage planning, goal thinning and validation, previews
preflight.sh, arm_cam_tune.shRead-only site checks; camera view and arm stow-pose tuning
dynamic_masking/The masking method: chunk_fusion_masker.py (fusion, object policy, geo gate), dynamic_object_mask.py (YOLO), dynamic_fusion.py + flowseek_flow.py (motion), motion_compensation.py
slam_integration/fusion_solver.py hooks the masker into the SLAM server, adds the live view and 3D carving; small patches for the server and camera gateway; idempotent install.sh
slam_tools/Server and gateway launchers, live view, replay, robot control (kachaka_ctl.py), camera check, cleanup, bird's-eye comparison
docs/Guides, validation reports and this page

Hardware and software

  • Robot: Kachaka mobile base (driven through the Kachaka API, move_to_pose), with a robot arm that holds the camera in a fixed stow pose during mapping.
  • Camera: Intel RealSense D435 RGB-D on USB 3.
  • Compute: NVIDIA Jetson (aarch64) running two Docker containers: a camera-to-ROS 2 gateway and a GPU container with the SLAM server and the masker.
  • Models: YOLOv9e-seg (semantic), FlowSeek (optical flow), the lab's ma-long SLAM backend; SAM 3.1 for offline evaluation.

Limitations and next steps

Known limitations

  • With the geo gate on, an object the detector cannot label is no longer removed by motion alone (DYNAMIC_GEO_GATE=none restores it).
  • A few thousand person points remain per run, from depth spill at a person's edges and legs behind furniture.
  • In blurred frames while the robot turns, the bridge can still paste a person shape onto a static object.
  • People far away or behind frosted glass are placed at the wrong depth by the model.
  • The held-rule thresholds come from only a few recordings.
  • Tinted glass lets depth see through it, which gives blotchy areas; this isn't a masking error.
  • A 0.40 m wall clearance removes much of the free space, so always check the preview.

Being evaluated

  • Giving YOLO its expected BGR colour order: on run2 it cut frames with a missed person from 11 to 3 (offline).
  • Stopping the bridge in blurred turning frames.
  • Using the robot's LiDAR pose as a prior to reduce SLAM drift.