A useful soft skills assessment starts with a decision, not a scoring form. Educators must define the behavior they need to observe, create an appropriate opportunity to demonstrate it, and collect enough evidence for the intended use.
This guide presents a seven-step framework for assessing soft skills in healthcare education. It covers observable behaviors, evidence sources, behavioral anchors, sampling, judgment, and feedback without pretending that professional judgment disappears.
For an introduction to why communication, empathy, teamwork, and professionalism matter—and how they can be taught—start with Michaella Masters’ overview of soft skill training in medical education. This article begins where that overview ends: designing the assessment system.
Soft skills, non-technical skills, or interpersonal skills?
The terms overlap, but they are not always interchangeable. “Soft skills” is accessible and familiar. Healthcare research often uses “non-technical skills” when discussing behaviors that support safe performance, while “interpersonal skills” usually narrows the focus to communication and relationships.
| Term | Typical scope | Examples |
|---|---|---|
| Soft skills | Broad educational and workplace term | Communication, empathy, teamwork, professionalism |
| Non-technical skills | Cognitive and social skills supporting safe clinical performance | Situation awareness, decision-making, teamwork, leadership |
| Interpersonal skills | Interaction with patients, families, and colleagues | Listening, explanation, rapport, conflict management |
A systematic review summarized by AHRQ PSNet identified 76 observer-based tools across healthcare settings and found substantial variation in their validation and usability. The practical implication is not to find a universal soft-skills score. It is to build a coherent local assessment around a defined purpose.
The worked example throughout this guide is closed-loop communication during a deteriorating-patient simulation. The learner should state a clear request, direct it to a named team member, confirm that the message was received, and follow up when the loop remains open.
| Blueprint component | Worked example |
|---|---|
| Capability | Team communication |
| Observable behavior | Gives a clear request, names the receiver, and confirms receipt |
| Context | Deteriorating-patient simulation |
| Evidence | Direct observation plus time-specific video markers |
| Raters | Facilitator and trained peer observer |
| Stakes | Formative |
| Next action | Repeat the scenario with a changed team and competing demand |
Step 1. Define the Decision Before Choosing an Assessment Tool
First, state exactly what the assessment will inform. A feedback conversation needs different evidence from a progression decision.
Write the purpose in one sentence. Then name the learner group, decision maker, setting, and consequences.
Record this decision in the blueprint. Otherwise, later tool choices may quietly change what the program intends to judge.
For example, a formative purpose might identify communication behaviors for the learner’s next simulated consultation. A higher stakes purpose might inform progression after several clinical placements.
Van der Vleuten and colleagues describe assessment as a continuum of stakes. Their programmatic model treats most observations as information rich data points. However, higher stakes decisions require more evidence and stronger safeguards.
The model is conceptual, so it does not prove that one program design produces better outcomes. It also allows an exception for mastery tasks.
Next, decide how educators will use the result. The NBME explains that frequency alone does not make assessment formative. Educators must use the evidence for feedback, instructional decisions, and continued learning.
Therefore, do not label an assessment formative when its scores simply enter a database. Plan the response before selecting the tool.
Videolab’s guide to formative and summative feedback can help teams clarify this first decision.
Step 2. Translate Professional Standards Into Observable Behaviors
Convert each selected capability into actions that an assessor can see or hear. Avoid rating broad impressions such as professionalism or empathy.
Start with the capability framework
Use the curriculum or regulatory framework that already governs the program. Then select only the capabilities relevant to the assessment purpose.
The GMC Generic Professional Capabilities framework includes professional values, professional skills, leadership, and team working. It translates professional responsibilities into educational outcomes for postgraduate curricula.
However, the framework does not provide a local rubric or observation count. Your team must still specify the performance evidence.
Write indicators that describe actions
Replace “communicates effectively” with indicators tied to a defined encounter. An indicator might require a clear handover, active information seeking, or confirmation of shared understanding.
The AHRQ TeamSTEPPS Team Performance Observation Tool shows this translation clearly. It organizes observable actions under team structure, communication, leadership, situation monitoring, and mutual support.
Still, TeamSTEPPS addresses teamwork rather than every soft skills construct. Use its structure as an example, not as automatic validation for another rubric.
Finally, remove indicators that require mind reading. Assessors can describe whether a learner checked understanding, but they cannot directly score private intentions.
Test each indicator against a short question: could two observers point to the same action in a recording?
The guide to measuring communication skills in healthcare provides examples of observation tools and complementary evidence sources for communication-focused assessments.
Step 3. Choose Evidence Sources That Match the Behavior
Choose each method because it reveals a relevant behavior in a suitable context. Do not add methods merely to make the program look comprehensive.
Match each method to its purpose
Direct observation can capture workplace interaction. Meanwhile, simulation can create comparable conditions for a specific challenge.
Video can preserve an encounter for later review. Multisource feedback can add perspectives from people who observe the learner in different relationships.
| Evidence source | Best suited to | Main limitation |
|---|---|---|
| Direct observation | Workplace behavior in context | Cases and observation opportunities vary |
| Simulation | Comparable, deliberately designed challenges | Transfer to practice cannot be assumed |
| Video review | Traceable behavior-level evidence and feedback | Requires governance and reviewer time |
| Multisource feedback | Patterns across professional relationships | Raters may observe different behaviors |
| Patient or standardized-patient feedback | Experienced communication and respect | Does not capture every professional capability |
Van der Vleuten and Schuwirth argue that complex competence needs integrated evidence across methods and occasions. However, they also warn that simply combining measures can compound error.
Use multisource feedback selectively
Lockyer and Sargeant describe multisource feedback as a four stage formative process. It collects observable workplace behavior, aggregates responses, reports patterns, and supports facilitated action planning.
Therefore, invite only raters who can observe the target behavior. Also, do not use multisource scores alone for progression decisions.
The paper is a narrative overview, not a primary study. Much of its evidence concerns practicing physicians rather than undergraduate learners.
Use video as evidence, not proof
In one formative OSCE study, Junod Perron and colleagues compared video based and direct feedback. Video sessions addressed more communication, professionalism, and reasoning content. Learners also participated more actively.
However, the single site formats differed in timing and duration. The study did not test later clinical performance.
For practical context, see Videolab’s guide to video in competency based medical education.
Create a simple evidence map. For each capability, list the method, context, observer, and decision that will use it.
Step 4. Build Behavior Anchors, Then Test Them Locally
Write rating anchors that distinguish visible levels of performance. Then test whether local assessors interpret those anchors consistently.
Anchor ratings to evidence
Start with examples at the weakest and strongest relevant levels. Next, describe the actions that separate those levels.
Avoid labels such as poor, adequate, or excellent without behavioral detail. Otherwise, each assessor must invent a private standard.
Also, connect every rating to a comment or recorded moment. This requirement makes the judgment easier to review and discuss.
For closed-loop communication, a three-level set of anchors could look like this:
| Level | Observable evidence |
|---|---|
| Needs development | Gives information or a request without naming the receiver or confirming receipt |
| Developing | Names the receiver but confirms receipt only after prompting or inconsistently |
| Meets the standard | Gives a clear request, names the receiver, confirms receipt, and follows up when no response occurs |
These anchors are an example, not a validated scale. A program should test them against representative performances and revise ambiguous language before use.
Pilot before using scores
Behavior anchors improve clarity, but they do not guarantee agreement. Holland and colleagues tested a nine point BARS across four nontechnical domains.
Three faculty rated 30 medical students in pediatric emergency simulations. Interrater agreement remained limited across situational awareness, decision making, communication, and teamwork.
However, this small study came from one institution and used three raters. Many learners lacked emergency simulation experience, and ratings covered a restricted range.
Therefore, the results do not show that BARS always fails. They show why every program should pilot its scale with representative cases.
Ask assessors to explain disagreements during the pilot. Then revise ambiguous indicators, instructions, examples, or scale points before consequential use.
Use several pilot cases, including borderline examples. Clear anchors must help assessors explain both agreement and disagreement.
Step 5. Plan Enough Observations, Contexts, and Rater Perspectives
Build a sampling plan that states when, where, how often, and by whom each behavior will receive observation.
Sample across relevant conditions
Performance changes with the case, team, patient, and setting. Therefore, one polished encounter cannot represent a learner’s usual practice.
Van der Vleuten and Schuwirth identify sampling as a major determinant of reliability. Their review emphasizes content, cases, examiners, and patients rather than standardization alone.
However, their testing time estimates come from different studies and designs. Do not turn those illustrative figures into universal requirements.
Instead, map each target behavior against the conditions that could change its interpretation. Then collect enough contrasting observations to test the emerging pattern.
Choose raters by access to behavior
Assign each rater only the behaviors that person can observe. For example, a patient and a team member see different parts of communication.
Lockyer and Sargeant report context dependent rater ranges for multisource feedback. However, they warn that weak observers can undermine credibility.
Consequently, more ratings do not always provide better evidence. The relationship, observation opportunity, and confidentiality plan also matter.
Finally, treat learner self assessment as a reflection input. Programmatic assessment guidance advises against using it alone for a decision.
A sampling table can reveal gaps early. Rows can list capabilities, while columns record settings, occasions, and observer roles.
Step 6. Aggregate Evidence Before Making Consequential Decisions
Define how reviewers will combine evidence before the first difficult case reaches them. This preparation limits improvised rules and unexplained averaging.
Keep observations traceable
Store the behavior, context, assessor role, date, and supporting comment together. If video exists, link the rating to the relevant moment.
Next, group evidence by capability rather than by assessment form. This view lets reviewers examine communication across placements, simulations, and different observers.
Van der Vleuten and colleagues recommend more information and safeguards as stakes rise. They also recommend scrutiny when evidence conflicts.
However, their model does not prescribe a universal observation count. Nor does it prove that committee review removes bias.
Use structured group judgment
For consequential decisions, give reviewers a shared evidence summary and explicit decision criteria. Also, document the reasons for the final judgment.
The ACGME Clinical Competency Committee Guidebook offers practical guidance on committee preparation, group process, documentation, feedback, and follow up. Its model comes from graduate medical education, so undergraduate programs must adapt it carefully.
When evidence conflicts, reviewers should examine context before averaging scores. They may need another observation, a different perspective, or clearer documentation.
Finally, define an appeal route for higher stakes decisions. Transparency matters because professional judgment remains part of the process.
Set a review trigger for missing or contradictory evidence. This rule prevents reviewers from improvising when uncertainty becomes uncomfortable.
Step 7. Return Specific Feedback and Audit the Assessment
Turn the evidence into a next action, then check whether the assessment still serves its stated purpose.
Close the feedback loop
Start with the observed behavior and its context. Then explain the effect, invite the learner’s interpretation, and agree on one next action.
The NBME treats formative assessment as an ongoing process, not a test type. Evidence must shape feedback, instruction, or subsequent learning.
Similarly, multisource feedback needs facilitated interpretation and action planning. A dashboard alone does not complete the process.
Video can make feedback more specific because both people can review the same moment. However, the Junod Perron study supports richer discussion, not proven performance transfer.
Videolab’s guides to giving feedback and implementing feedback models can help teams turn observed evidence into an actionable conversation.
Review the assessment system
Audit missing observations, score distributions, rater patterns, learner responses, and decision delays. Also, ask whether the process remains feasible for staff.
The AHRQ TeamSTEPPS measurement guidance links measurement planning to the intervention’s aims. However, its framework addresses teamwork implementation rather than every healthcare education assessment.
Schedule the first audit before launch. Early review gives the team a planned point for revising forms, workflows, and guidance.
Use the framework as a connected assessment system
The seven steps form one sequence. Define the decision, translate the capability into behavior, choose suitable evidence, develop and pilot anchors, sample across relevant conditions, combine evidence transparently, and return feedback that changes the next performance.
Videolab can support this workflow by capturing or uploading encounters and connecting comments to specific moments. Educators can review the same evidence against agreed criteria and compare performance across attempts. The platform supports traceability and feedback, but it does not replace sampling, assessor judgment, or program governance.
Contact Videolab to discuss how recorded encounters and time-specific feedback can support a soft-skills assessment program.
