A camera-and-ejector clod sorter on a 1 m potato feed belt fails in a small number of recognisable ways: ejections land a fixed distance off, ejections scatter, the wrong finger fires, or clods never trigger at all. Each pattern points at a different subsystem. Read the pattern first, then work the checks below in order. Cheap checks (arithmetic, saved images, latency logs) come before expensive ones (cameras, valves, ejector fabrication).
Design basis from the application: stationary line, 1 m wide feed belt, standard belt speed 0.42 m/s on a geared motor (speed adjustable if a VFD or motor change is fitted), objects up to about 70 mm, target 20-30 t/h, 20 pneumatic fingers across the width, 1-2 global-shutter cameras, an encoder, and a PLC or real-time controller driving the valves. Missed clods and occasional rejected potatoes are acceptable. Commercial optical sorters and a steel-roller clod separator already exist; the roller loses reliability on wet, soft clay clods straight after harvest, and water is excluded because moisture raises storage-rot risk.
Read the misfire pattern before touching hardware
Log where every ejection lands relative to its target. The pattern tells you which check to run.
| What you see | Most likely cause | Go to |
|---|---|---|
| Every ejection lands the same distance early or late along the belt | Uncompensated fixed delay: wrong camera-to-ejector distance, missing actuation lead, or wrong encoder scaling | Measure fixed delay and jitter separately |
| Ejections scatter along the belt; worse when belt speed or load rises | Jitter: arrival-time triggering from the vision PC, PLC scan time, valve or cylinder response spread, product slip | Latch the encoder count at exposure |
| Timing is right but the neighbouring finger fires, or clods are clipped at the edge | Lateral calibration error or belt wander across the 50 mm finger pitch | Size the 20-finger ejector bank |
| Clods pass with no ejection command issued | Classification miss, clod hidden under potatoes (layer too thick), or poor lighting contrast | Calculate belt loading; prove classification on saved images |
| Bulk potato rejection | Wet potato skin classified as clay, threshold too aggressive, or lighting changed | Prove classification on saved images |
| Accurate on a slow belt, drifts when belt speed changes | Delays stored in milliseconds instead of encoder counts, or actuation lead not scaled with speed | Latch the encoder count at exposure |
| Accurate at start of shift, degrades later | Dirty lens or light window, LED output drift, wet belt, different soil batch | Set exposure, resolution and frame rate |
First check: is the error constant or scattered? A constant offset is a calibration fix in minutes. Scatter is an architecture problem, and no parameter tweak cures it.
Rule out the mechanical separators once, then stop revisiting them
Two low-tech routes are already closed for this line:
- Steel-roller rebound separator. It works when clods are dry and hard, such as at grading before shipment. Directly after harvest, wet soft clay clods rebound like potatoes, so the divider cannot separate the two streams reliably.
- Water (flotation, washing, soaking clods apart). Excluded. Seed potatoes stay in storage for months, and added moisture raises bacterial infection and rot risk while slowing drying.
Both facts justify vision. They also give you a reuse decision: the traditional separator's frame, 1 m input belt and two discharge paths are a workable mechanical base once the roller is removed. Confirm the belt drive and head-roller geometry allow a finger bank at the discharge before you commit to that base.
Calculate belt loading before you pick a camera
A camera sees only the top layer. A clod under a potato is invisible, so single-layer flow decides whether classification can work at all. Run the numbers first.
Coverage check, with labeled assumptions: treat each object as a circle in plan view, 70 mm diameter for the 95 objects/m case (0.00385 m2 each, about 37% of a 1 m2 belt section) and about 50 mm for the 190 objects/m case (0.00196 m2 each, also about 37%). Roughly a third of the belt area is covered at 0.42 m/s at full rate. Objects will touch and some will overlap, and a real feed is not evenly spaced. Random feed clumping matters more than the average.
Decision:
- Loading at the target rate and 0.42 m/s is 19.8 kg/m2. If a photo of the belt at full feed shows stacked objects, you have two levers: raise belt speed (halves loading each time you double speed) or throttle the infeed. The source of the potatoes is controllable, so throttling is legitimate.
- Higher belt speed thins the layer but shrinks the timing budget in ms per mm and changes the trajectory of objects leaving the belt end. Speed and layer thickness trade against each other; set them together.
- A singulating or spreading section ahead of the camera (a bar, a short accelerating belt, or a vibrating spreader) is a mechanical fix that costs less than a second camera.
The narrow-test-rig question resolves here: per-object timing depends on belt speed, not width. Width multiplies the object rate and the number of cameras and ejector channels. It does not make the timing problem easier.
Prove the classification on saved images before buying ejectors
Vision is the most tractable part with off-the-shelf libraries and SDKs, but wet clay and potato skin look alike under poor lighting. Test that on a bench, with one camera and a fixed light, before spending on pneumatics.
- Collect images of real wet, soft clods and real potatoes from the current harvest, on a moving belt or a hand-pulled belt section, under the enclosure lighting you intend to use.
- Label them by hand: potato, clod, and touching pairs. Keep a separate hold-out set from a different day or field. Soil type and moisture change the appearance.
- Try the cheap discriminators in order: colour and brightness in a diffuse light; texture (potato skin is smooth, clods are rough and irregular); shape; then a small trained detector or segmentation model. Trained models learn texture and shape cues that fixed thresholds miss.
- If wet clay and skin still overlap, test lighting variants before changing software: diffuse dome versus low-angle light, cross-polarised light to suppress the specular sheen of wet surfaces, and a near-infrared channel. Water and soil respond differently in the NIR, but this must be tested on your samples.
- Score the result as a confusion matrix: clods missed, potatoes wrongly flagged. Your tolerance is loose (some missed clods, some rejected potatoes). Define numeric limits for both before you start, so you know when it is good enough.
If saved images cannot be separated by eye at full resolution, no ejector timing will fix it. Stop here and fix lighting or the feed.
Set exposure, resolution and frame rate from belt speed
Three numbers size the camera side. All three follow from belt speed and the 70 mm maximum object size.
- Resolution. mm per pixel = field width / horizontal pixels. Example with a labeled assumption: two cameras, each covering 500 mm plus overlap, on a sensor with about 2000 pixels across gives roughly 0.25 mm per pixel. That is far finer than needed to classify 50-70 mm objects, so a modest sensor may already be enough. Pick the sensor from the classification test, not from the catalogue.
- Frame rate and coverage. Trigger the camera by encoder distance, not by a free-running clock. Set the trigger step = field length along the belt - 70 mm, so every object is fully visible in at least one frame. Illustration, assuming a 300 mm field length along the belt: step = 230 mm, which at 0.42 m/s is about 1.8 frames per second per camera. Convert each detection to belt coordinates (encoder count at capture plus position in the image) so an object seen in two overlapping frames is counted once.
With two cameras across the 1 m width, overlap the fields by at least 70 mm, the maximum object size, so no object is cut by a camera boundary. Alternatively, a line-scan camera with encoder-triggered lines removes frame stitching for a constant-speed belt, at the cost of a different camera and lighting setup. Put the lens and lights behind a window and keep the window clean; dust and moisture on the optics show up as classification drift over a shift.
Measure fixed delay and jitter separately
Fixed delay is calibrated out. Jitter is what ruins the sort. The source question was which contributes most: encoder accuracy, pneumatic response variation, PLC latency, or belt slip. In practice the order to check is:
- Product slip and tumbling between camera and ejector. An object that rolls or slides after imaging goes where the software does not expect. No timing fix helps. Check with dummy targets and a short camera-to-ejector distance; slope, belt cleats or a short, flat, dry belt reduce it.
- How the vision result is timestamped. If the PLC fires from the time a message arrives from the vision PC, every operating-system and network delay becomes position error. Carry the encoder count captured at exposure inside the message (next section).
- Valve and cylinder response spread. Command-to-contact time varies with supply pressure, hose length, cylinder load and temperature. Measure it.
- Encoder resolution. Rarely the limit. See the formula below; 1 mm per count is already fine against 50 mm finger pitch.
Measure valve response: mount an end-of-stroke sensor (or a proximity switch at the finger tip), command the finger 100 or more times at production pressure, and log command-to-sensor time. The mean becomes t_act in the fire equation. The spread (max minus min) times belt speed is the position error you cannot calibrate out.
The +/-10 mm figure is a design choice: about a seventh of the largest object. Pick your own tolerance from the smallest clod you want to hit. Add the spreads from the list above; if the sum exceeds the allowed spread, slow the belt or fix the largest contributor.
Latch the encoder count at exposure and fire on counts
The architecture in the source (camera, vision computer, encoder, PLC valve timing) is correct. The rule that makes it work: nothing in the timing path may depend on when a message arrives.
mm_per_count = roller_circumference_mm / (encoder_PPR * 4) # x4 quadrature
at exposure: N_cap = encoder count latched by hardware (camera strobe or trigger output
into a PLC high-speed input, or the encoder-triggered acquisition itself)
per object: x_mm = position across belt; y_mm = position along belt from image edge
finger = floor(x_mm / 50) # 1000 mm / 20 fingers = 50 mm pitch
(also fire finger+1 if the object spans a pitch boundary)
N_fire = N_cap + (D_ej - y_mm - v*t_act) / mm_per_count
D_ej = calibrated distance camera-edge to ejection line
t_act = measured mean command-to-contact time
v = current belt speed from the encoder
PLC: push (finger, N_fire, hold_time) into a FIFO;
when N_now >= N_fire, pulse the finger output for hold_time
Rules for the encoder and the drive:
- Mount the encoder on a non-driven roller or a measuring wheel that rides the belt. A driven pulley slips on a wet, dirty belt, and the encoder then reports motion the product did not make.
- Because the lead term uses
v * t_act, calculate it from measured speed. That keeps timing correct when the VFD changes belt speed. - Give the vision PC the encoder count via the PLC or a shared counter, not the PC's own clock. If message latency varies, N_cap still identifies exactly where the object was.
- Put the FIFO, comparison and output in the PLC or real-time controller; keep only detection and classification on the PC.
Size the 20-finger ejector bank
Twenty fingers over 1 m gives a 50 mm pitch. Objects up to 70 mm span at least two adjacent fingers in the worst case, so the fire logic must handle multi-finger ejections and the mechanical design must allow adjacent fingers to fire together.
- Stroke and force. Design against the object mass. From the rate estimate above, average object mass is roughly 104-208 g; treat clods as similar in the absence of a measured value, and weigh a sample of wet clods to replace that assumption. A finger has to deflect that mass sideways or into the reject path within the time the object is at the ejection line.
- Air consumption. Free air per stroke = (pi/4 x bore^2 x stroke) x (p_abs / p_atm). Multiply by the strokes per second on a busy section to size the compressor and supply line. Do this with your cylinder data before buying valves.
- Valves. Choose valves whose datasheet response time supports the cycle rate and the timing spread you budgeted. Mount them close to the cylinders; shorter tubing reduces response time and its spread. Regulate pressure so it does not sag when many fingers fire.
- Alternatives. Air-blast nozzles are the standard ejector in commercial sorters, but 100-200 g clods need considerably more impulse than small products. Test with your heaviest wet clods before committing either way, and check whether an off-the-shelf ejector module fits the 1 m width.
- Wet clay. Clay builds up on fingers and guides and changes clearances. Make the finger tips easy to wipe or replace, and log the response time periodically.
- Belt tracking. Belt wander shifts the object-to-finger alignment across the width. A wander of half the 25 mm half-pitch is enough to hit the neighbouring finger, so set belt tracking and calibrate the camera's lateral scale against finger positions.
Decide what a 30 cm test rig buys you
Narrowing the belt to about 30 cm (or blocking off the rest of the 1 m belt with side guides) does three things: it needs one camera and about six fingers, it cuts the object rate roughly by two thirds, and it removes stitching between cameras. It does not change the timing physics per object, which depends on belt speed. That makes it the right way to find jitter, slip and classification problems cheaply.
Scale in this order: first low speed and low throughput, then raise belt speed, then raise the feed until the layer is as thick as production, then add width. At each step re-measure hit position scatter and the confusion matrix. If one step breaks accuracy, you know which variable did it.
Scope matters too. A warrantied, documented unit at full production rate is a team project by machine-builder standards (an estimate of 10-15 people and 40-50 weeks has been quoted for that). An experimental in-house unit that must work reliably on your own line is a different target, and its risk sits mostly in the timing chain and lighting, not in the mechanics.
Commission the resolving branch and verify the sort
Run this sequence once the classification test passes on saved images. Each step depends on the one before it.
- Fix the belt speed and measure it against the encoder over a known length of marked belt. Record
mm_per_count. - Measure the camera-to-ejection-line distance
D_ejphysically, then trim it with a test object. - Log finger command-to-contact time over 100 or more cycles at working pressure. Enter the mean as
t_act. Record the spread. - Check the controller scan or fast-task time and confirm the FIFO and output logic run on it.
- Calibrate lateral position: place a marked object over each finger position and confirm the vision x coordinate maps to the right finger.
- Dry run with dummy targets (dried clods or marked blocks) at low speed. Measure landing offset along and across the belt. A constant offset means adjust
D_ejort_act; scatter means return to the jitter list. - Repeat at the target belt speed and change the speed in steps. Offset must stay put; if it moves with speed, the lead term is not scaled by speed or a delay is stored in time rather than counts.
- Run real harvest product with the ejectors disabled but the classifier logging. Compare flagged objects against a hand count from the same belt section.
- Enable ejection. Sample both output streams over several loads and hand-sort them: clods left in the potato stream, potatoes in the reject stream. Compare with your numeric tolerances.
Keep these logs: encoder count at capture, object position and class, N_fire, actual finger sensor timestamp, and product flow rate. When performance changes, they show whether timing, classification or feed caused it.
FAQ
Can one or two global-shutter cameras cover a 1 m belt for clod detection?
Yes on resolution: at 0.42 m/s and 50-70 mm objects, even about 0.25 mm per pixel (two cameras, roughly 2000 pixels across each, illustrative) is far finer than needed. The limits are lighting contrast between wet clay and potato skin and a single-layer feed, so test both on saved images before choosing sensors.
Does a narrower test belt make the encoder-to-valve timing easier?
No. Timing error per object depends on belt speed (0.42 mm of travel per millisecond at 0.42 m/s), not belt width. A narrower belt reduces cameras, ejectors and object rate, which makes debugging cheaper, but the jitter budget stays the same.
Can I let the vision PC trigger the valves directly over the network?
Do not use message arrival time as the trigger. Latch the encoder count at exposure, send it with the detection, and have the PLC or real-time controller fire when the live count reaches the computed N_fire, so PC and network delay stop mattering.
When should I stop and ask for official support?
Stop when hit-position scatter stays above your tolerance after you have measured valve response spread, controller scan time and belt slip, or when classification accuracy does not improve with lighting changes and a hold-out set. Then send your logged encoder counts, latency measurements and image samples to the official support channel of the camera, PLC and valve manufacturers, or engage a commercial sorter builder for a hardware and lighting review.