VinUniversity

A 4-DOF arm with computer vision

A pick-and-place arm that finds red, blue and white blocks with an overhead camera and works out where they are. All the vision runs on an NVIDIA Jetson Nano, and an Arduino Uno drives the servos. I built the block detection, the camera calibration and the pixel-to-world transform.

Overhead camera view of the arm with its claw gripper beside three blocks, blue, red and white, each in a detection box with a centre dot and a label giving its colour and X, Y and Z position.
The overhead view: each block detected, with its centre and its position in the arm's coordinates.

Goal

Pick up small blocks wherever they lie on the work area, with one overhead camera and edge hardware. The camera has to say not only which colour a block is but where it is in the arm's own coordinates, accurately enough for the gripper to close on it.

What I built

Detection

Each frame from the 1080p, 60 FPS camera is converted to HSV, where colour is separate from brightness, and thresholded into red, blue and white masks: red and blue by hue, white by high brightness and low saturation, with the thresholds tuned by hand. A contour from a mask counts as a block only if its area lies between 500 and 10,000 pixels, its aspect ratio between 0.8 and 1.2, and it falls inside the work area. Each block then gets its centre and the colour of the mask at that point.

Camera calibration

A chessboard with 11 × 11 inner corners and 3 cm squares, photographed 44 times, gives the camera's intrinsics and lens distortion through OpenCV, with every corner refined to sub-pixel accuracy.

Pixel to world

Ten points measured by hand on the work area, and found in the image, give the camera's pose through solvePnP. With that pose and the intrinsics, a block's centre in the image becomes a position in the arm's coordinates, which the arm's inverse kinematics then turns into joint angles.

Result

Detection IoU ranged from 0.92 to 0.96, and F1 from 0.92 to 0.97, across the three colours and two block sizes (35 × 35 × 30 and 50 × 30 × 35 mm), under controlled lighting. Pick-and-place succeeded in 77–84% of trials, depending on colour and size. The ten reference points reprojected with a mean error of 10.2 pixels (standard deviation 5.4), about 0.5% of the image width, most likely from hand-measured points, leftover lens distortion and the small number of points.

What I learned

  • Picking a colour space where the property you threshold on is separate from the lighting.
  • Filtering detections by size, shape and position, which removes most false positives before anything smarter is needed.
  • Camera calibration end to end: intrinsics and distortion from a chessboard, then the camera's pose from known points.
  • Where reprojection error comes from, and that more and better-measured reference points are the cheapest fix.
  • Running a whole vision pipeline on a Jetson Nano at the edge.