How to Use DeepSeek's Vision Capability to Verify Hardware Kits Before They Ship
A practical guide to using DeepSeek's image input to check photos of kitted equipment against a packing list before a multi-site China rollout ships — a seven-step bench workflow with an example prompt, what a photo check cannot catch, and how to keep site data out of the frame.
Published
The short answer: When a kit for a store on the other side of the country ships with the wrong power cord or the 24-port switch instead of the 48-port one, the error is found by the one person on site least able to fix it. DeepSeek's API now accepts images alongside text, which makes it practical to photograph each kit on the staging bench and have the model compare what it sees against the packing list, line by line, before the box is sealed. It flags mismatches for a person to check; it cannot see inside sealed packaging, read every tiny label, or tell you whether anything works.
Picture a logistics coordinator in a staging warehouse in Suzhou, three weeks into a rollout for a retail chain opening 30 new stores across Sichuan, Chongqing and Yunnan. Each store gets the same kit: a PoE switch, two ceiling access points with brackets, a firewall appliance, a label printer, patch cables, and power cords for everything. Kits are packed from an ERP pick list exported to Excel, checked by eye against a printout, and sealed.
Store 14, in Chengdu, opens its box on a Saturday morning. The switch is the 24-port model, not the 48-port one the floor plan was designed around, and the firewall came with a UK-plug power cord that fits nothing on the wall. There is no IT person on site. Replacement parts go out by courier, and the installer booked for a two-hour visit needs a second trip.
Nothing about this is unusual. The pick list was right. The shelves held both switch models side by side, with part numbers that differ by two characters, and the checker had forty more kits to finish before the courier cut-off.
This article is about adding one cheap, fast step to that process: a photo of each opened kit, compared by a vision-capable AI model against the packing list before the box is sealed. It does not replace the human check; it gives it something specific to look at.
Why a Wrong Cable or Missing Adapter Only Gets Found After the Kit Arrives
Kitting errors are rarely caused by a bad list. They come from the physical step between the list and the box: similar-looking items stored side by side, substitutions made when a bin runs empty, a regional power-cord variant pulled from the wrong carton, an accessory bag never added because the main unit looked complete.
The manual check that is supposed to catch this has a structural weakness. The checker reads the list and looks for each item. A 24-port and a 48-port switch from the same product family are the same colour, often close to the same size, and carry their model numbers on a small sticker. A China-standard power cord and a UK-plug cord look identical until you look at the plug end, which is usually coiled at the bottom of the box.
For a single-site install, a mistake like this costs an afternoon. For a multi-site rollout into stores with no local IT staff, it costs a failed install visit, a courier round trip and a delayed opening. The error is cheap to catch in the warehouse and expensive everywhere else.
What a Vision-Capable Model Can Actually Check From a Photo
DeepSeek announced an experimental vision model, `deepseek-v4-flash-vision-exp`, on 21 August 2026. Within weeks, its API documentation described image input on the main `deepseek-flash` model: JPEG, PNG, GIF and WebP images, sent in user messages in the same OpenAI-compatible Chat Completions format as text, either as base64, as a URL, or uploaded through a Files API. Given that pace, check DeepSeek's current documentation for the model name, image limits and pricing before you build, rather than trusting any article, including this one.
For kitting, the useful capability is narrow and practical. Given a clear photo of an opened kit laid out on a bench and a text packing list, the model can describe what it sees, match items to list lines, and say which lines it cannot find evidence for.
Comparing a kitted-equipment photo against a packing list, item by item
The basic check is presence. You give the model the packing list for Store 14 — part number, description and quantity per line — and a top-down photo of the kit with every item laid out and nothing stacked. You ask it to go through the list in order and, for each line, say whether it can see the item and how many.
The value is not that the model looks better than a person. It is that it goes through every line every time, does not get tired at kit 38, and produces a written result that a person can check against the photo in under a minute. A line marked "not visible" is an instruction to go and look.
Flagging a plausible-but-wrong substitution (right brand, wrong model) a rushed human check misses
The more valuable check is the one a tired person skips: reading the label. If the packing list says `SW-48P-POE` and the switch in the photo has a front panel with visibly fewer ports, or its label reads `SW-24P-POE` in a close-up, the model can flag the mismatch even though the item looks right at a glance. The same applies to power cords laid out with plug ends facing the camera: a UK Type G plug looks clearly different from a China-standard one.
A Practical Workflow — From a Packing List to a Photo-Verified, Ship-Ready Kit
This is the version a logistics coordinator can set up in a week with the pick list they already export and a phone mounted over the staging bench. The part numbers are illustrative.
Step 1 — Standardise the packing list per site. Export the pick list for each kit from your ERP or Excel into a plain table: site code, line number, part number, description, quantity, and a "check point" column saying what must be visible — for example "plug end facing camera" for power cords, or "front panel ports visible" for switches.
Step 2 — Fix the photo setup. Mount a phone or a cheap document camera at a fixed height over a bench with a plain, contrasting mat. Mark a layout grid on the mat so items go in the same place every time: switch top-left, access points and brackets top-right, cords and adapters along the bottom with plug ends out.
Step 3 — Take two photos per kit. One wide shot of the full layout, and one close-up of the labels that matter most — the switch model sticker, the firewall label and the plug ends. Name the files with the site code and kit number.
Step 4 — Obscure what the model does not need. Before anything is sent, crop out or cover shipping labels with store addresses, and asset tags or serial numbers unless your process actually requires the model to read them.
Step 5 — Send the photos and the list together. A short script or internal tool sends both images plus the packing list for that site in one request. A prompt that works as a starting point:
"You are checking a hardware kit before shipment. Below is the packing list for site CD-014, followed by two photos of the kit laid out on a bench. For each packing-list line, in order, report: the line number, whether the item is visible (yes / no / unclear), the quantity you can count, and any visible difference from the listed part number or description, such as a different model label, port count or plug type. Do not guess. If a label is too small or blurred to read, say 'unreadable' rather than inferring. Then list any item in the photos that matches no line. End with one line: PASS if every line is 'yes' with matching quantity and no differences, otherwise REVIEW."
Step 6 — A person resolves every REVIEW. The checker compares the flagged lines with the physical kit, fixes the error or overrides the flag with a short note, and only then seals the box. A PASS still gets a quick human glance.
Step 7 — Keep the record. Store the photos, the model's output and the checker's decision against the site code. When a site reports a problem, you can see exactly what left the warehouse, and over time which items need a better photo angle.
AI-Assisted Photo Verification vs a Manual Checklist Inspection vs No Formal Kitting QA
- AI-assisted photo verification: every line checked in order on every kit, with a written, timestamped record and specific flags for a person to resolve. It adds a camera setup, a small per-kit API cost and a data-governance decision about what the photos contain.
- Manual checklist inspection: no new tools, and an experienced checker can open a box and sense that something is wrong. Reliability drops with volume and fatigue, similar-looking substitutions pass, and there is usually no record beyond a tick on paper.
- No formal kitting QA: fastest and cheapest in the warehouse. Every error is found by the site, which on a multi-site rollout means by people without the parts or knowledge to fix it, at the most expensive point in the process.
Where a Photo Check Isn't Enough
A photo shows presence and appearance. It cannot tell you that a switch powers on, that its firmware is correct, or that an access point is not dead on arrival. Functional testing — powering each device and confirming it takes the staged configuration — is a separate step, and a photo check does not replace it.
It also cannot see inside sealed packaging. If ten patch cables are bagged together, or an access point's mounting screws sit in a sealed pouch inside the retail box, the model can at best confirm the bag or box is there. Open and lay out what needs counting, or accept that those lines rely on the supplier's packing, not the photo.
Vision models can also miscount small identical items and misread small or angled model-number labels. That is why the prompt asks for "unclear" and "unreadable" rather than a best guess, and why every REVIEW goes to a person. Treat the model's PASS as "nothing obviously wrong", not as a guarantee.
At scale — hundreds of kits a week, serialised assets — a photo check is one control inside a QA process that still needs scan-based pick confirmation, serial capture, functional testing and sign-off.
Getting This Right — Verification Accuracy, Site-Data Confidentiality, and When to Bring in IT
Do not quote an accuracy figure for this to anyone, including yourself. Measure it on your own kits: for the first few weeks, record every case where the model flagged something that was fine, and every error a site found that the model passed. That is the only accuracy number that means anything for your process.
Decide what goes into the photos before the first one is sent. DeepSeek's hosted API is operated from China, and kit photos can easily include store addresses on shipping labels, asset tags, device serial numbers and staging network details written on stickers. Crop or cover what the check does not need, keep paperwork out of the layout, and check your company's policy on sending operational images to an external AI service.
Handle the API key like any other production credential: kept in a secret store rather than in the script, and rotated when the person who set it up moves on.
This is where outside help earns its place. Building the image pipeline, the logging and the key handling is the kind of work AI support covers, and the devices in the kit still need someone to configure and support them once they arrive, which is ordinary IT support work. The paperwork side of the same shipment is covered in our guide to DeepSeek for China hardware import customs documentation, and matching what shipped against your asset register in DeepSeek for IT asset inventory reconciliation.
If you would rather not run kitting and staging in-house, Brocent — founded in Beijing in 2007 — can stage, check and ship equipment kits for multi-site rollouts through its warehouse-as-a-service offering.
FAQ
How accurate is a vision model at spotting a wrong item?
There is no single honest figure, and any number quoted without your items, your photos and your lighting should be ignored. Accuracy depends heavily on the photo: a flat, well-lit layout with labels facing the camera helps; a jumbled box does not. Measure it on your own kits for a few weeks before relying on it.
Does this replace a human QA check entirely?
No. It changes what the person does: instead of re-reading the whole list for every kit, the checker resolves the specific lines the model flagged and glances over the passes. A person still decides whether the box ships.
Can it check functional condition, not just presence?
No. A photo cannot show whether a device powers on, holds its configuration or has a fault. Functional testing stays a separate step.
Does this work for kits with dozens of small components?
Partly. Vision models can miscount small identical items such as screws, cable ties or bagged patch leads. Pack small parts into labelled bags counted at packing, and use the photo for the larger, distinguishable items.
Is it safe to send kit photos to DeepSeek's API?
That depends on what the photos contain and on your company's policy. The hosted API is operated from China. Keep addresses, serials and asset tags out of frame unless the check needs them, and get the policy question answered before the rollout, not after.
Which DeepSeek model should we use?
Whichever model DeepSeek's current documentation lists as accepting images. Check the Vision guide for the current model name, limits and pricing when you build.
Share:
Ready to take action?
Turn these insights into a roadmap for your business.
Book a 15-minute no-obligation consultation with our APAC IT experts. We'll review your current setup and provide a tailored IT roadmap within 24 hours.
Free Checklist
10 Critical Checks Before Expanding IT to Greater China
PIPL compliance, network segmentation, bilingual helpdesk setup, and more — everything your IT team needs before Day 1 in China.
Request the checklist →📬 Monthly Asia IT Insights
China compliance updates, cybersecurity alerts, and IT tips for APAC teams — once a month.
No spam. Unsubscribe anytime.