Run the Hardest-SKU Test
Vision-guided picking has improved enormously, and the honest framing is still that item handling is unsolved in general and solved for specific ranges. The question is whether your range is one of them.
The Bot Scout hardest-SKU test is deliberately unfair to the vendor: take the twenty items your human pickers complain about most, and ask for the success and intervention rate on those. A pick rate averaged across easy SKUs describes a demonstration; the difficult tail describes your operation.
The items that defeat picking systems are predictable: transparent and reflective packaging, deformable bags, tangled or nested items, very light objects, and anything where the grip point depends on how it happens to be lying.
| Item type | Difficulty | Why | What to require from a vendor |
|---|---|---|---|
| Rigid boxed goods | Low | Predictable faces and grip points | Rate at your case sizes |
| Polybagged items | Medium | Deformable and variable shape | Success rate on your bag types |
| Transparent or reflective packaging | High | Perception struggles to resolve them | A live test with your items |
| Nested or tangled items | High | Grip separates more than one item | Double-pick rate, not just success rate |
| Very light or very small items | High | Vacuum and vision both marginal | Evidence at your SKU dimensions |
The Exception Path Is the Design
Every picking deployment has items the robot refuses. The system design question is what happens to those: whether they route to a human station automatically, how quickly, and whether the order can still ship complete and on time.
A deployment that handles 90 percent of picks and dumps the remaining 10 percent unpredictably into a human queue can be harder to staff than manual picking, which is the utilisation problem described in the warehouse automation guide.
Fleet coordination and mobility are separate purchases from the picking cell itself; compare those in the autonomous mobile robots guide, and read internal-fleet coverage of Amazon Robotics in our guide with the caution that those systems are not sold externally.
- Test on your twenty hardest SKUs, not a representative sample.
- Ask for intervention rate and double-pick rate, not just success rate.
- Design the exception route to a human station explicitly.
- Confirm how order completeness is maintained when picks fail.
- Re-test after any significant change to the SKU range.
Perception Tuning Is an Ongoing Cost
Picking performance depends on models tuned to the item range. A new supplier, a packaging change, or a seasonal range shift can degrade performance without anything mechanical changing.
Ask who retunes the system, how quickly, and at what cost. A vendor-only tuning model means every packaging change becomes a support ticket with a lead time attached.
Safety in a shared space still applies. OSHA notes that many robot incidents occur during setup, testing, and maintenance, which for a picking cell means the hours spent clearing jams and adjusting grippers.
Bottom Line
Warehouse picking robots are evaluated on the hard tail of your SKU range and on the exception path for what they refuse. Require intervention rates on your own items before comparing pick rates.
Send your twenty hardest SKUs to every vendor and ask for intervention rate, not pick rate.
FAQs
What items are hardest for picking robots?
Transparent or reflective packaging, deformable polybags, nested or tangled items, and very light or very small objects where the grip point depends on how the item is lying.
What is intervention rate and why does it matter?
The share of picks needing a human. It determines staffing far more than headline pick rate, because the exceptions arrive unpredictably and still have to ship on time.
How should a picking robot be piloted?
With your own hardest SKUs rather than a representative sample, measuring success rate, intervention rate, and double-pick rate, and testing the exception route to a human station.
Does picking performance change over time?
Yes. New suppliers, packaging changes, and seasonal range shifts can degrade performance. Confirm who retunes the perception models, how fast, and at what cost.