A 4-DOF arm with computer vision
A pick-and-place arm that finds red, blue and white blocks with an overhead camera and works out where they are. All the vision runs on an NVIDIA Jetson Nano, and an Arduino Uno drives the servos. I built the block detection, the camera calibration and the pixel-to-world transform.
Goal
Pick up small blocks wherever they lie on the work area, with one overhead camera and edge hardware. The camera has to say not only which colour a block is but where it is in the arm's own coordinates, accurately enough for the gripper to close on it.
What I built
Detection
Each frame from the 1080p, 60 FPS camera is converted to HSV, where colour is separate from brightness, and thresholded into red, blue and white masks: red and blue by hue, white by high brightness and low saturation, with the thresholds tuned by hand. A contour from a mask counts as a block only if its area lies between 500 and 10,000 pixels, its aspect ratio between 0.8 and 1.2, and it falls inside the work area. Each block then gets its centre and the colour of the mask at that point.
Camera calibration
A chessboard with 11 × 11 inner corners and 3 cm squares, photographed 44 times, gives the camera's intrinsics and lens distortion through OpenCV, with every corner refined to sub-pixel accuracy.
Pixel to world
Ten points measured by hand on the work area, and found in the image, give the camera's pose through solvePnP. With that pose and the intrinsics, a block's centre in the image becomes a position in the arm's coordinates, which the arm's inverse kinematics then turns into joint angles.
Result
Detection IoU ranged from 0.92 to 0.96, and F1 from 0.92 to 0.97, across the three colours and two block sizes (35 × 35 × 30 and 50 × 30 × 35 mm), under controlled lighting. Pick-and-place succeeded in 77–84% of trials, depending on colour and size. The ten reference points reprojected with a mean error of 10.2 pixels (standard deviation 5.4), about 0.5% of the image width, most likely from hand-measured points, leftover lens distortion and the small number of points.
What I learned
- Picking a colour space where the property you threshold on is separate from the lighting.
- Filtering detections by size, shape and position, which removes most false positives before anything smarter is needed.
- Camera calibration end to end: intrinsics and distortion from a chessboard, then the camera's pose from known points.
- Where reprojection error comes from, and that more and better-measured reference points are the cheapest fix.
- Running a whole vision pipeline on a Jetson Nano at the edge.