Methodology

How assessment and verification work

This page explains what we measure, how scores are produced, where AI is used, where a human decides, and how any result can be challenged. It is written to be read by the people being assessed, not only by the people hiring.

What we measure

We publish a multidimensional scorecard rather than a single mysterious number. Each dimension is scored separately, carries its own confidence value, and links to the evidence behind it.

DimensionWhat it meansEvidence
Role CompetencyKnowledge and applied understanding of the target role.Knowledge items and scenario tasks.
PracticalQuality of a real work simulation.Work artifact scored against a rubric, plus reviewer check.
CommunicationWritten and voice client communication.Situational responses and intro video transcript.
EnglishFunctional workplace English.Written, listening and speaking tasks.
AI FluencyEffective, safe and verified AI-assisted workflow.Scenario and practical exercise.
Portfolio EvidenceStrength and verifiability of prior work.Source files, version history, client confirmation.
ReputationPerformance inside Taskopedia.Reviews, completion, repeat work, disputes.
ReliabilityDelivery behaviour, not personality inference.On-time milestones, responsiveness, attendance where contracted.

How a score is calculated

A dimension score is the weighted mean of the scored items within it. Confidence is calculated separately, as the product of how much of the blueprint the evidence actually covers, how reliable the assessment instrument is, and whether a human reviewed the result.

dimension_score = Σ(item_score × item_weight) / Σ(item_weight)
confidence      = evidence_coverage × assessment_reliability × review_factor
published_score = round(dimension_score)

The weights are held in versioned configuration that is reviewed and published, never inside a prompt. A result is always pinned to the assessment version it was taken under, so changing a rubric never silently rewrites scores that already exist.

Confidence is deliberately kept separate from the score. A high score with low confidence means we do not yet have enough evidence, not that the person is weak. Where confidence falls below threshold, the result goes to a human reviewer instead of being published as a precise number.

Role weightings in use

Different roles weight the dimensions differently. These are starting configurations and are being validated against real job outcomes; they will change by version as we learn which signals actually predict performance.

brand designer

brand-designer-v1.0

  • Practical35%
  • Role competency20%
  • Portfolio evidence15%
  • Communication10%
  • AI fluency10%
  • Reputation/reliability10%

customer success

customer-success-v1.0

  • Communication25%
  • Role competency20%
  • Practical scenario20%
  • English15%
  • Reputation/reliability10%
  • Tools/AI fluency10%

full stack developer

full-stack-developer-v1.0

  • Practical code35%
  • Role competency25%
  • Debugging/security15%
  • Communication10%
  • AI coding fluency10%
  • Reputation5%

digital marketer

digital-marketer-v1.0-derived

  • Practical campaign brief30%
  • Channel fundamentals and analytics25%
  • Copy judgement and reporting20%
  • AI research and verification15%
  • Reputation/reliability10%

virtual assistant

virtual-assistant-v1.0-derived

  • Prioritisation scenario25%
  • Organisation, scheduling and tools25%
  • Written communication and client updates20%
  • English10%
  • AI productivity workflow10%
  • Reputation/reliability10%

sales sdr

sales-sdr-v1.0-derived

  • Outreach and pitch25%
  • Prospecting, ICP reasoning and CRM25%
  • Objection handling20%
  • English10%
  • Ethical AI-assisted research10%
  • Reputation/reliability10%

ai engineer

ai-engineer-v1.0-derived

  • Practical modelling task35%
  • Model, data and evaluation knowledge25%
  • Explaining a result15%
  • Responsible AI workflow15%
  • Reputation/reliability10%

data analyst

data-analyst-v1.0-derived

  • Practical analysis30%
  • SQL, statistics and reporting30%
  • Turning analysis into a recommendation20%
  • AI-assisted analysis10%
  • Reputation/reliability10%

cybersecurity analyst

cybersecurity-analyst-v1.0-derived

  • Threat, vulnerability and response knowledge35%
  • Practical triage30%
  • Incident reporting15%
  • AI-assisted detection10%
  • Reputation/reliability10%

devops engineer

devops-engineer-v1.0-derived

  • Practical pipeline task35%
  • Cloud, CI/CD and observability30%
  • Runbooks and handover15%
  • AI-assisted operations10%
  • Reputation/reliability10%

qa engineer

qa-engineer-v1.0-derived

  • Practical test suite30%
  • Test design and tooling30%
  • Defect reporting20%
  • AI-assisted testing10%
  • Reputation/reliability10%

ux designer

ux-designer-v1.0-derived

  • Practical flow or screen30%
  • Research, patterns and accessibility20%
  • Portfolio evidence15%
  • Design rationale15%
  • AI-assisted design10%
  • Reputation/reliability10%

video editor

video-editor-v1.0-derived

  • Practical edit35%
  • Editing, motion and delivery craft20%
  • Portfolio evidence15%
  • Responding to a brief10%
  • AI-assisted production10%
  • Reputation/reliability10%

content writer

content-writer-v1.0-derived

  • Practical writing piece30%
  • Brand voice and editing to feedback20%
  • Research, structure and search intent20%
  • English10%
  • AI-assisted drafting and fact-checking10%
  • Reputation/reliability10%

finance accountant

finance-accountant-v1.0-derived

  • Bookkeeping, reconciliation and reporting35%
  • Practical month-end task30%
  • Explaining the numbers15%
  • AI-assisted finance workflow10%
  • Reputation/reliability10%

hr recruiter

hr-recruiter-v1.0-derived

  • Sourcing, screening and people process30%
  • Practical shortlist exercise25%
  • Candidate and manager communication20%
  • English5%
  • AI-assisted sourcing10%
  • Reputation/reliability10%

ecommerce specialist

ecommerce-specialist-v1.0-derived

  • Storefront, merchandising and analytics30%
  • Practical store task30%
  • Customer and supplier messaging20%
  • AI-assisted listing and support10%
  • Reputation/reliability10%

legal support

legal-support-v1.0-derived

  • Documents, research and confidentiality30%
  • Practical drafting task25%
  • Client correspondence15%
  • English15%
  • AI-assisted review5%
  • Reputation/reliability10%

cad technician

cad-technician-v1.0-derived

  • Practical drawing task35%
  • Standards, tooling and take-off30%
  • Mark-ups and revisions15%
  • AI-assisted drafting10%
  • Reputation/reliability10%

medical back office

medical-back-office-v1.0-derived

  • Terminology, coding and records35%
  • Practical claim or record task25%
  • Patient and practice correspondence15%
  • English10%
  • AI-assisted administration5%
  • Reputation/reliability10%

research analyst

research-analyst-v1.0-derived

  • Practical research brief30%
  • Design, sources and data collection25%
  • Writing up findings20%
  • AI-assisted research and verification15%
  • Reputation/reliability10%

supply chain coordinator

supply-chain-coordinator-v1.0-derived

  • Sourcing, planning and documentation35%
  • Practical purchasing exercise25%
  • Supplier negotiation and chasing20%
  • AI-assisted planning10%
  • Reputation/reliability10%

project manager

project-manager-v1.0-derived

  • Practical plan or recovery exercise30%
  • Planning, risk and delivery method25%
  • Stakeholder updates25%
  • AI-assisted coordination10%
  • Reputation/reliability10%

automation specialist

automation-specialist-v1.0-derived

  • Practical automation build35%
  • Process mapping, tooling and integration30%
  • Handover documentation15%
  • AI inside a process, with checks10%
  • Reputation/reliability10%

game developer

game-developer-v1.0-derived

  • Practical prototype35%
  • Engine, gameplay and performance30%
  • Working with designers and artists15%
  • AI-assisted tooling10%
  • Reputation/reliability10%

embedded engineer

embedded-engineer-v1.0-derived

  • Practical firmware task35%
  • Firmware, interfaces and connectivity30%
  • Technical documentation15%
  • AI-assisted debugging10%
  • Reputation/reliability10%

Where AI is used, and where it is not

AI assists with

Structuring an imported CV into a draft profile, which the talent must confirm.

Generating question variants from an approved competency blueprint.

Applying a written rubric to a text, voice or artifact response.

Transcribing audio and generating captions.

Turning an employer description into a structured role brief.

Phrasing a match explanation from features that already exist in the ranking data.

AI never

Decides the competency weights or the pass threshold at run time.

Publishes a score without schema validation and version metadata.

Issues a verification badge or removes access to opportunities on its own.

Adds a reason to a match explanation that is not present in the feature data.

Infers personality, honesty or emotion from a face or a voice.

Claims to determine whether a design or document was AI-generated.

Every evaluator run records the provider, model, prompt template version, rubric version and a confidence value, so any published score can be reconstructed and audited later.

Verification states

StateMeaningWhat employers see
Profile createdSelf-declared information, not yet checked.No badge.
Identity verifiedIdentity check passed.Identity badge only.
Assessment completedRequired assessments done, quality review may still be pending.Scores provisional and visible to the talent only.
Taskopedia VerifiedRole thresholds met, required checks passed, human QA complete.Verified badge with date and assessment version.
Verified, active evidenceRecent assessment or work evidence on file.Full badge.
Verification staleEvidence has aged beyond the freshness window.Badge annotated as reassessment due, or downgraded by policy.
Under reviewA material flag or appeal is open.Affected scores limited until resolved. This does not imply wrongdoing.

Common questions

Does AI decide whether I pass an assessment?
No. AI applies a written rubric to produce a per-item score, but the weights, the pass thresholds and the aggregation are held in versioned configuration and computed in code. Low-confidence and borderline results are routed to a human reviewer before anything is published, and no badge is issued without human quality review.
What does the confidence value next to a score mean?
Confidence is calculated separately from the score, as the product of how much of the assessment blueprint the evidence actually covers, how reliable the instrument is, and whether a human reviewed the result. A high score with low confidence means we do not yet have enough evidence, not that the person performed poorly.
Can I challenge a result I think is wrong?
Yes. Any score that materially affects your verification status or your access to opportunities can be appealed. An appeal creates a tracked case with a named reviewer who was not involved in the original decision. If a score is overridden, the original result is preserved alongside the new one with the rationale and policy basis recorded.
Do you assess personality, confidence or honesty from my video?
No. We do not use facial expression, emotion inference, honesty detection, attractiveness scoring or accent ranking. Communication is scored on observable criteria such as clarity, structure, professionalism, completeness and expectation setting, using the transcript and a written rubric.
Can I use AI in a practical assessment task?
Yes, and we assess how well you do it. AI fluency is a scored dimension covering whether you use AI effectively, verify its output and understand its limitations. What matters is that you disclose how AI was involved and that the resulting work is good. We do not claim to detect AI-generated work from the artifact alone.
Will my scores change if you update a rubric?
No. Every result is pinned to the assessment version it was taken under, and published assessment versions are immutable. Changing a rubric affects future sessions only. If you want a result under a newer version, you take a reassessment.
How long does verification stay valid?
Each role track has a freshness window. Once your most recent assessment or completed work falls outside it, your profile is marked as due for reassessment rather than presented to employers as current evidence.

Integrity and appeals

Assessments use randomised item pools, server-side timing and per-section limits. Where integrity signals are recorded, they are disclosed in advance, and they never automatically fail anyone. A flag routes a session to a human reviewer who looks at the actual evidence.

We do not use invasive webcam proctoring by default. Any future use would require a clear justification, explicit consent and legal review first.

Any score that affects verification or access to opportunities can be appealed. An appeal creates a tracked case with a named reviewer. If a score is overridden, the original result is preserved alongside the new one with the reviewer, rationale and policy basis recorded.