A note on language
The downloadable document and spreadsheet are in Spanish, like the rest of the templates here. This page explains what they contain and the reasoning behind each rule, so you can judge whether they fit before translating.
When to use it
During the session, with the prototype or the product in front of you. It is what the moderator and the observer keep open while someone tries to get something done.
It is not the conversation script. That is the interview guide, which already covers the guided walkthrough. This is the layer the guide does not have: what was expected of each task, whether it happened, what it cost, and what was observed.
It is not the plan either. Who gets recruited, with what method and for what decision is agreed beforehand, in the research plan.
It works for moderated testing, in person or remote. Unmoderated testing has tasks and metrics but no observation, so half the sheet stays empty and you are better off going straight to the spreadsheet.
Before you use it
The success criterion is written before you test. That is the rule everything else rests on. With no prior criterion, "the task failed" is an opinion to be argued about in the results meeting; with one, it is a fact you can check. And a criterion written after seeing the data is always met.
Separate the scenario-level goal from the site-level one. "Finds the form in under 2 minutes" is measured on one task; "90 % complete the purchase unaided" is measured on the product. Mixing them produces goals nobody can check.
"With help" is not success. It is its own category. Without it, a task the moderator rescued with a hint gets recorded as completed, and the number is inflated exactly where the problem was.
Time is only recorded if the task was completed. Averaging the time of failed tasks means nothing: someone who gives up after 20 seconds is "faster" than someone who succeeds in three minutes.
The moderator does not take notes. Or takes very few. If you are writing, you are not watching. This sheet is meant to be filled in by the observer, with the moderator marking only what cannot wait.
Everything in brackets gets replaced.
The document
One sheet per task, repeated as many times as the session has tasks. It is printed or filled in on screen, one copy per participant. It is in Spanish.
{/* --- INICIO DOCUMENTO --- */}
Proyecto: [cliente · producto o flujo] Participante: [ID, nunca el nombre] · Perfil: [según el screener] Modera: [nombre] · Observa: [nombre] Fecha: [DD/MM/AAAA] · Duración prevista: [60] min Qué se pone delante: [prototipo en baja / en alta / producto en producción] Modalidad: [presencial · remoto] · [moderado · no moderado] · Herramienta: [cuál]
Las tres últimas líneas no se deciden acá: vienen del plan de investigación. Se copian en la ficha porque quien observa necesita tenerlas a la vista, y porque son las que explican, al releer, por qué un resultado dice lo que dice.
Metas medibles del testeo
Se completan antes de la primera sesión. Son las mismas que van en la hoja Metas de la planilla.
| Categoría | Nivel | Tarea | Prioridad | Requisito medible |
|---|---|---|---|---|
| Tiempo | Escenario | [tarea] | [Alta/Media/Baja] | [El 90 % encuentra X en menos de 3 minutos] |
| Precisión | Escenario | [tarea] | [ ] | [El 90 % llega a X con menos de 2 clics equivocados] |
| Éxito | Sitio | [tarea principal] | [ ] | [El 90 % completa la tarea sin ayuda] |
| Satisfacción | Sitio | — | [ ] | [El 80 % puntúa 4 o más en una escala de 1 a 5] |
Antes de empezar
- Consentimiento firmado y recibido
- Grabación autorizada y andando
- Prototipo abierto y probado en el dispositivo de la sesión
- Plan B listo si el ambiente de prueba se cae
Lo que se dice al abrir: el guion está en la guía de entrevistas. Lo esencial: no hay respuestas correctas, evaluamos el producto y no a la persona, y puede detenerse cuando quiera.
Tarea [n] · [nombre de la tarea]
Escenario que se lee en voz alta
[Una situación, no una instrucción. «Imagina que quieres X y entras al sitio…», nunca «haz clic en el botón azul».]
| Meta | [qué debería lograr la persona] |
| Criterio de éxito | [qué tiene que pasar para contarla como completada, escrito antes] |
| Punto de partida | [pantalla o estado desde donde empieza] |
| Datos que necesita | [lo que hay que tener a mano: correo, número, archivo] |
| Ruta esperada | [los pasos que el equipo supone que va a seguir] |
| Suposición que estamos probando | [la creencia del equipo que esta tarea pone a prueba] |
Registro
| Completó | Sí · Con ayuda · No |
| Tiempo (s) | [solo si completó] |
| Errores | [n de acciones equivocadas antes de retomar] |
| Dónde se detuvo | [pantalla o paso exacto] |
| Qué dijo ahí | [cita textual, no parafraseada] |
| Qué hizo con el cuerpo | [dudas, suspiros, volver atrás, releer, buscar el mouse] |
Notas de quien observa
[Lo que no cabe en las casillas de arriba.]
Después de las tareas
Las preguntas van abiertas primero y con escala después, para que el número no ancle lo que cuenta.
- ¿Cómo sentiste la experiencia en general?
- ¿Cuál fue la tarea más difícil? ¿Qué la hizo difícil?
- ¿Hubo algún momento en que no supiste qué hacer? Cuéntame ese momento.
- Si pudieras cambiar una sola cosa, ¿cuál sería?
- ¿Hay algo que esperabas encontrar y no estaba?
- En una escala de 1 a 5, donde 5 es lo mejor, ¿cómo calificarías la experiencia? · [ ]
- ¿Por qué ese número y no uno más?
La pregunta 7 es la que rescata la escala. Un 4 sin explicación no dice nada; un 4 con «porque el paso del correo me dio desconfianza» es un hallazgo.
Cierre de la sesión
Impedimentos que enfrentó la persona: [ ]
Preguntas que hizo durante el testeo: [lo que preguntó es lo que el producto no le respondió]
Preguntas que hizo al final: [ ]
Lo que se llevó quien observa, en una frase: [ ]
Nota de método y limitaciones
Con cinco o seis participantes esto describe lo que pasó en esas sesiones, no proporciones de la población: «4 de 5» no es «80 % de los usuarios».
De los tres indicadores, la eficacia es el más sólido. El tiempo depende mucho del dispositivo y del contexto de la sesión, y la satisfacción declarada al final de una sesión moderada tiende a ser más alta que la real, porque quien participa le está hablando a una persona.
Los hallazgos que salgan de acá se priorizan aparte, en la matriz de severidad. Esta ficha registra; no decide qué se arregla primero.
Caduca. Revisar en [fecha]: un protocolo escrito para un prototipo que ya cambió mide otra cosa.
{/* --- FIN DOCUMENTO --- */}
The results spreadsheet
The sheets above are filled in by hand, one per participant. This spreadsheet is what turns them into numbers: it downloads separately, has five tabs and calculates on its own.
{/* --- INICIO RESULTADOS --- */}
Qué trae la planilla
Metas — las cuatro categorías de meta medible, con su nivel y su umbral. El requisito se escribe en prosa para poder leerlo, y además en número —tiempo máximo y éxito mínimo— porque una planilla no puede comparar contra una frase.
| Categoría | Nivel | Qué mide |
|---|---|---|
| Tiempo | Escenario | Cuánto cuesta llegar |
| Precisión | Escenario | Cuántos desvíos hubo en el camino |
| Éxito | Sitio | Si se completa la tarea principal |
| Satisfacción | Sitio | Cómo quedó la persona al final |
Participantes — una fila por persona: ID, perfil y su puntaje de satisfacción de 1 a 5, validado para que no entre un 7.
Sesiones — una fila por participante y tarea. Es lo que trae el protocolo llenado a mano: si completó, cuánto demoró, cuántos errores, y qué pasó. La columna «cumple la meta» cruza el tiempo con el umbral de esa tarea, y dice Sin meta cuando no hay umbral escrito, en vez de inventar un veredicto.
Resumen — se recalcula solo y excluye las filas de ejemplo:
| Indicador | Qué es | De dónde sale |
|---|---|---|
| Eficacia | % de tareas completadas con éxito | Sesiones · columna D |
| Eficiencia | Tiempo y errores promedio por tarea | Sesiones · columnas E y F |
| Satisfacción | Promedio de la escala 1 a 5 | Participantes · columna D |
Más el desglose por tarea, que es el rango desde el que se arma el gráfico.
Cómo se usa
- Escribe las metas antes de la primera sesión.
- Una fila por participante, con su satisfacción al cerrar.
- Una fila por participante y tarea, transcribiendo las fichas.
- Borra las filas marcadas
EJantes de compartir. El resumen ya las excluye del cálculo.
{/* --- FIN RESULTADOS --- */}
Check before you use it
- Are the goals written before the first session, or were they written while looking at the results?
- Does every task have a success criterion that does not depend on the watcher's judgement?
- Do the scenarios describe a situation, or do they tell the person where to click?
- Is "with help" kept separate from "yes"?
- Are the moderator and the observer different people?
- Does the participant ID replace the name everywhere on the sheet?
- Did you write down verbatim quotes, or summaries of what you hoped to hear?
- Did you delete the
EJrows from the spreadsheet before sharing it?
Grounding
The measurable-goals structure — the time, accuracy, success and satisfaction categories, and the distinction between scenario level and site level — is adapted from the Measurable Usability Goals template by Usability.gov (U.S. Department of Health and Human Services), a U.S. government work and therefore in the public domain. The examples and the rest of the document are my own.
The per-task sheet and the closing questions come from a protocol I use on retail projects; the five session stages, from a usability-testing workshop I put together for a product team in 2022.
What surrounds this template: the research plan agrees the study, the screener recruits, the informed consent authorises the recording, the interview guide is the session script, and the severity matrix prioritises whatever comes out of this.
The deliverable page is Usability Report.