A secure messenger pilot in 30 days: what actually needs to be measured
A pilot does not prove that people like an interface. It must show whether identities, devices, revocation, critical workflows and recovery hold under controlled friction.
A defensible 30-day pilot defines one critical workflow, measurable success, permitted test data and stop conditions before the first sign-in. It then evaluates identity, devices, roles, disruption and recovery separately — and ends with continue, correct or stop.
- Success and stop criteria are written before the first sign-in.
- Identity, device change, revocation and recovery are tested alongside messaging.
- Expectation, observation and impact are documented separately.
- The pilot ends with an evidenced decision, not an average rating.
Day 0: a claim that may be disproved
The pilot starts with one sentence: ‘Team X can complete workflow Y under condition Z within N minutes without external help.’ This deliberately names the people, action, pressure and measurement — and permits an unambiguous failure.
Product interest, click counts or ‘it feels secure’ are not pilot objectives. A useful objective answers a decision: should the workflow advance, be corrected or stop for now?
- Accountable owner and two to ten active testers
- One core workflow with start, finish and abort point
- Permitted data class and expressly excluded data
- Metric, target and stop condition before first sign-in
Gate 1: identity and device are distinct controls
A successful sign-in proves very little by itself. The pilot must establish which person acts through which device and session. First access, a second device, device loss, session revocation and return are therefore evaluated separately.
BSI guidance calls for unique user identifiers, need-based permissions and managed grant, change and withdrawal. NIST Zero Trust adds that network location or ownership must not create implicit trust; user and device are authenticated and authorised before access.
- Can one session be revoked without removing every device?
- Is a removed or lost device effectively excluded?
- Are role, device and security-relevant sign-in traceable?
- Does recovery work without an informal bypass?
Gate 2: measure the critical workflow as a timeline
Instead of touching twenty features, test one workflow deeply: accept an invitation, open the Mission Room, understand context, verify a message or file, record a decision and confirm completion. Measure time, help, failed attempts and outcome for every step.
Speed alone is not the objective. A fast but misunderstood workflow is worse in a security context than a slightly slower controlled one. Comprehension and outcome quality belong on the same timeline.
- Median and slowest successful run, not average alone
- Share completed without help and share with failed attempts
- Correct result and correctly understood context
- Break point under time pressure, interruption or role change
Gate 3: controlled friction, not a happy path
Weeks two and three introduce deliberate friction rather than sabotage: weak connectivity, app restart, a second device, revoked permission, an outdated file or an interrupted workflow. Each exercise has an expected behaviour and a safe recovery path.
BSI emergency-management guidance calls for regular exercises, documented results and evaluation for improvement. A messenger pilot does not simulate a catastrophe; it exercises foreseeable disruption in a controlled way.
- Connection loss during a critical step
- Role or membership change during the workflow
- Lost device and targeted session revocation
- Recovery with unambiguous, intact context
The VENTEX Evidence Matrix
Every finding is split into expectation, observation, impact, reproducibility and decision. This stops a feeling from being rewritten as a technical cause and stops an isolated error from being generalised without context.
One-to-five scores support trend analysis but never replace the finding. One reproducible blocker may matter more than twenty positive convenience scores. The VENTEX pilot board therefore combines severity, task outcome and qualitative observation.
- Expectation: what should the workflow have done?
- Observation: what visibly happened?
- Impact: convenience issue, friction or security-relevant blocker?
- Decision: accept, correct, retest or stop?
Stop criteria protect more than a perfect score
Stop criteria are agreed in advance because time and success pressure otherwise turn boundary breaches into exceptions. The pilot pauses for ambiguous identity, ineffective revocation, data loss, uncontrolled permission or use of excluded data.
Stopping is not project failure. It is a successful control point: contain the environment, understand the cause and restart only after a documented correction.
- Ambiguous identity or unauthorised access
- Revocation does not reliably remove access
- Loss, mixing or misattribution of test data
- A blocker cannot be safely bounded or reproduced
The 30-day cadence
Day 0 bounds team, data, workflow and target. Week one measures access, devices and the undisturbed core workflow. Week two introduces device and role change. Week three tests interruption and recovery. Week four closes blockers and repeats decisive measurements.
NIST’s SSDF is outcome-focused and encourages risk-based adaptation and continuous improvement. ENISA’s maturity framework follows a compatible logic: assess the baseline, identify gaps and prioritise improvement. The schedule serves the decision, not the calendar.
- Day 0: boundaries, roles, data and stop criteria
- Days 1–7: access, devices and undisturbed workflow
- Days 8–21: friction, revocation, interruption and recovery
- Days 22–30: close blockers, repeat measurement, decide
Day 30: continue, correct or stop
The final record fits on one page: objective, sample, data class, measured workflow, results, open blockers, security boundaries and next decision. Screenshots and raw feedback belong in the evidence, not a marketing slide.
Continue means only that the next bounded step is defensible. Correct names an owner, deadline and repeat measurement. Stop records why the present use case is not responsible. None of these decisions is a certification or blanket production approval.
FAQ
A pilot does not prove that people like an interface. It must show whether identities, devices, revocation, critical workflows and recovery hold under controlled friction.
For one bounded core workflow, two to ten active testers are often sufficient. Distinct roles and devices matter more than a large uncontrolled sample.
Not in the controlled VENTEX pilot without separate explicit approval. The programme is designed for synthetic, anonymised or non-sensitive internal test data.
No. It produces defensible product and process observations but does not replace independent security review, certification or regulatory approval.
The proportion of testers who correctly complete the pre-defined core workflow without help. Read it with failed attempts, comprehension and unresolved blockers.
