How to Organise MLPS (等保) Audit Evidence with DeepSeek
MLPS assessments stall on evidence, not controls. A working method for inventorying documents, mapping them to the 等保 requirement list with DeepSeek, handling the bilingual load, and surfacing gaps early enough to fix them.
Published
The short answer: MLPS assessments rarely fail on the security controls themselves — they stall because nobody can find, organise and translate the evidence that proves the controls exist. DeepSeek is well suited to that middle layer: mapping scattered documents to a requirement list, drafting bilingual descriptions, and flagging gaps early. It must never be the thing that decides whether a control is adequate.
Most foreign-invested companies going through their first MLPS (等级保护, "graded protection") assessment discover the same thing about six weeks in. The security controls are not missing. Firewalls exist, backups run, accounts are managed, logs are retained. The problem is that proving it means producing a specific document, in a specific form, in Chinese, cross-referenced to a specific requirement clause — and that document currently lives as a screenshot in someone's WeChat history, a policy PDF written in English by a European parent company, and an unwritten practice three engineers have in their heads.
This article is about that gap, and nothing else. It does not explain what MLPS or PIPL require — the site already carries a full MLPS and PIPL compliance checklist for foreign companies in China for that. What follows is the evidence workflow: inventory what you have, map it against what is being asked, draft what is missing, package it — using DeepSeek to absorb the volume and the bilingual load that make this job slow.
Why Evidence Collection, Not Compliance, Is What Actually Delays an Assessment
An MLPS assessment has a shape that surprises people who have been through ISO 27001 or SOC 2. The assessor works from a structured requirement list at your system's assigned grade, and for each item wants something concrete: a policy, a configuration screenshot, a log excerpt, a signed record, a system report. The interview and the on-site inspection matter, but the document pack does most of the work.
The delay comes from four places, none of them security-related.
The evidence is scattered by default. Nobody creates a compliance evidence folder in advance. Change approvals are in a ticketing system, the account review is an annual email thread, backup verification is a monthly screenshot in a group chat, the network diagram is in whichever engineer's Visio file is most recent. All of it exists. None is where a document pack expects it.
The language does not match. A foreign-invested enterprise usually inherits group-level policies in English. The assessment is conducted in Chinese, against Chinese-language standards, by assessors who read what you hand them. Somebody has to translate — not just linguistically, but into the vocabulary the standard uses, so a clause about "privileged access management" is recognisably answering the requirement it is meant to answer.
Nobody knows what is genuinely missing until late. Teams work down the requirement list, find a real gap at item 140 of 200, and only then need to build and operate a control long enough to have evidence of it working. Found in week two, that is a manageable remediation. Found in week nine, it moves the assessment date.
The person doing it has another job. Evidence preparation is almost never anyone's full-time role. It lands on an IT manager also running the estate, or a compliance owner covering several jurisdictions, in fragments of days.
AI does not fix the fourth problem, but it compresses the first three — and by compressing them, surfaces the third early, which is where most of the value is.
Turning a Requirement List Into an Evidence Register
The single artifact this exercise produces is an evidence register: one row per requirement, saying what document answers it, where it is, what state it is in, and who owns it. Build the register first and let it drive the work, rather than collecting documents and hoping they cover the list.
Mapping what you have to what is asked for, and naming the gaps early
Start by getting the requirement list into machine-readable form — a spreadsheet with one row per control item, carrying the clause reference and the requirement text. Your assessor or consultancy normally supplies this; if they supply a PDF, converting it is the first task rather than a reason to work inside the PDF.
Then produce a plain inventory of what you actually hold. Not a curated list — a raw one. Policies, network diagrams, procedures, configuration exports, screenshots, ticket exports, training records, vendor contracts, previous audit reports. Filename, location, language, date, one-line description. A few hundred rows is normal.
The mapping step is where DeepSeek earns its place. Given the requirement text and the inventory, it can propose which documents plausibly answer which requirement, and — more usefully — mark the requirements where nothing in the inventory looks like an answer. A prompt that works well in practice is deliberately conservative:
"Here is a list of MLPS control requirements, and a list of documents we hold with descriptions. For each requirement, list the documents that could plausibly serve as evidence, ranked by how directly they address it. Where nothing in the inventory plausibly answers a requirement, say so explicitly and do not suggest a weak match. Do not assess whether the control is adequate — only whether a document exists that speaks to it."
That last instruction matters more than it looks. Without it, models reach for a plausible-sounding match rather than reporting a gap, and a false match is worse than a blank cell — it makes the register look complete when it is not.
The output will be wrong in places, which is fine: correcting a proposed mapping is far faster than building one from an empty spreadsheet, and the gap list it produces in week one is what protects your assessment date.
The bilingual problem — evidence in English, assessment conducted in Chinese
This is where DeepSeek beats a general-purpose Western model, for a mundane rather than ideological reason: the source material is bilingual and the target vocabulary is Chinese regulatory terminology. Models trained primarily on Chinese-language corpora handle 等保 vocabulary and the conventions of Chinese compliance documents more naturally, with less of the stiffness that makes a translated policy read as translated. Three jobs here are worth handing over:
- Translating existing policy documents into Chinese that uses the standard's own terminology, rather than a literal rendering. A group information-security policy translated word-for-word can be technically accurate and still fail to visibly answer the clause it is meant to answer.
- Drafting Chinese-language descriptions of controls that exist but were never written down. The control is real; the document is not. Dictate what actually happens, and have the model produce a first-draft procedure in the register's expected form.
- Producing bilingual summaries for the parent company. A parallel English summary of each Chinese document costs almost nothing and prevents a late objection from a group CISO who cannot read the pack.
For all three, the model produces a draft. A named human — ideally someone who will be in the room during the assessment — signs off every document before it enters the pack.
A Practical Workflow — Inventory, Map, Draft, Human-Review, Package
Run it in this order — the order is the method, the tool is interchangeable.
Inventory (2–4 days). Build the raw document list above. Resist the urge to tidy or judge as you go — a document you dismissed as irrelevant is often the only evidence of a control.
Map (1–2 days). Produce the first-pass register with the model, then walk it manually, checking two things: that proposed matches are real, and that the gap list is honest. Send the gap list to management the same week — it is the exercise's highest-value output.
Draft (1–2 weeks, overlapping remediation). Where the control exists and only the paperwork is missing, draft the documents. Where the control itself is missing, that is a remediation project — no amount of drafting substitutes for implementing it and running it long enough to generate evidence.
Human-review (continuous). Every generated or translated document gets read end to end by someone who knows whether it is true. A generated procedure describes what a reasonable organisation would do, not necessarily what yours does — and an assessor who finds a procedure that contradicts the interview has found a much bigger problem than a missing document.
Package (2–3 days). Assemble in the structure the assessor expects: indexed by clause, register as the cover map, consistent naming, dates and version numbers. Dull, and it materially affects how the assessment goes — an assessor who finds things quickly asks fewer questions.
AI-Assisted Evidence Preparation vs a Compliance Consultant vs Doing It in Spreadsheets
- AI-assisted preparation is fastest at the volume work — mapping hundreds of documents to hundreds of requirements, translating, drafting first versions — and it is available the day you decide to start. It has no standing in the assessment, no relationship with the assessor, and no opinion worth trusting on whether a control passes. Its output is raw material, not a submission.
- A compliance consultant brings the judgement AI cannot: what this assessor in this city actually expects, whether a control as implemented will be accepted, and how much a given gap matters. That is where the real risk sits, and it is worth paying for — which is exactly why you do not want those hours going on document sorting and translation.
- Spreadsheets and human effort alone is most companies' default, and it works — for a small system with a short requirement list and someone who has done it before. At larger scope it becomes a serialised bottleneck through one person's attention, surfaces gaps late, and drops the bilingual load on whoever happens to be bilingual.
The combination that works: AI for volume, spreadsheets as the durable register, consultant hours concentrated on judgement calls and the assessor relationship rather than document handling.
What AI Must Not Decide
Two questions have to be routed to a qualified assessor or China counsel every time, and no prompt makes them safe to answer internally.
"Is this control adequate for our grade?" Adequacy is a judgement made against the standard as applied by a licensed assessment body, with local practice baked in. A model can tell you a document appears to address a clause. It cannot tell you the implementation will be accepted, and a confident wrong answer here is worse than no answer, because it stops you asking.
"Is this gap material?" Whether a shortfall blocks certification, gets a corrective-action period, or is noted and passed is not a technical determination. Route it to the assessor or to counsel, and get the answer in writing before you plan around it. The same rule applies to PIPL obligations for the data your systems hold.
Getting This Right — Evidence Confidentiality, Data Residency, and When to Bring in IT
An MLPS evidence pack is one of the most sensitive document collections your company will ever assemble: network topology, security configurations, account management practices, known gaps, remediation plans — a précis of how to attack you, indexed for convenience. Decide these three things before the first document goes into a chat window.
Which deployment you are using. A public consumer chat interface, an API tier with commercial data-handling terms, and a privately deployed open-weight model on your own infrastructure are three different risk positions. DeepSeek publishes open-weight models that can be self-hosted, which for evidence work is a genuinely relevant option — worth evaluating properly rather than defaulting to the consumer app because it is what someone already had open. Check current terms directly with the provider; commercial data-handling policies change.
What actually needs to go in. Most of the value comes from document descriptions, requirement text and policy prose — not raw configuration exports, credential stores or live logs. Redact aggressively: if full content is not needed to map or translate a document, do not paste it.
Where the output lives. The register and the drafts are as sensitive as the sources. They belong in your managed document store under the same access controls as the evidence, not a personal drive or a chat history.
Choosing the deployment model, writing data-handling rules your team can actually follow, and building the prompts and register templates so this is repeatable next cycle is AI+ Support work. The controls themselves — and the monthly gap analysis and configuration backup discipline that make an evidence pack assemblable rather than reconstructable under pressure — sit with managed IT security services, where an Asia-based security control centre and certified consultants do this as standing practice. The estate work underneath — patching, account reviews, backup verification, which generate the evidence as a by-product of being run properly — is managed IT support. If an assessment date is set and the evidence position is unclear, talk to us now rather than in week nine.
Frequently Asked Questions
Can we put security evidence through an AI tool at all?
That depends on your deployment and internal policy, not on a general rule — a self-hosted open-weight model is a very different position from a consumer chat app. Decide the deployment, redact what does not need to be sent, and write the rule down. The failure mode is not a considered decision to use AI; it is an engineer pasting a firewall configuration into a public chat window because nobody said not to.
Does the evidence need to be in Chinese?
Assessments are conducted in Chinese against Chinese-language standards, and evidence an assessor can read directly generates fewer questions. Practice on accepting English group policies varies, and your assessor is the authority on what they will accept — ask early and specifically rather than assuming either way.
Who decides whether a control passes?
A licensed assessment body, against the standard — not your team, not a consultant, and not the model. Everything you generate internally is preparation for that judgment call, never a substitute for it.
How early should evidence collection start?
Earlier than feels necessary. The binding constraint is rarely the paperwork — a genuine control gap found late needs the control implemented *and* operated long enough to produce evidence it works. Running the inventory and mapping pass in the first fortnight converts a late surprise into an early one.
How accurate is AI mapping of documents to requirements?
Good enough to be a first draft, not good enough to submit. Expect to correct a meaningful share of proposed matches. The gap list is generally more reliable than the matches — convenient, because the gap list is the part with schedule consequences.
How does this interact with our PIPL obligations?
They are separate regimes that overlap in practice, and how they interact for your systems and data is a legal question. Our MLPS and PIPL compliance checklist covers what each requires; whether your evidence and data handling satisfy both is something to confirm with China counsel.
Can the same register be reused for the next assessment?
Yes, and that is most of the return on doing it properly. A maintained register — updated as controls change, with owners attached — turns the next cycle from a reconstruction project into an update. The companies that find their second assessment easy kept the register alive rather than filing it.
Share:
Ready to take action?
Turn these insights into a roadmap for your business.
Book a 15-minute no-obligation consultation with our APAC IT experts. We'll review your current setup and provide a tailored IT roadmap within 24 hours.
Free Checklist
10 Critical Checks Before Expanding IT to Greater China
PIPL compliance, network segmentation, bilingual helpdesk setup, and more — everything your IT team needs before Day 1 in China.
Request the checklist →📬 Monthly Asia IT Insights
China compliance updates, cybersecurity alerts, and IT tips for APAC teams — once a month.
No spam. Unsubscribe anytime.