You are not your user.
Testing with one is better than none.
Evaluation is integral to the design process. Evaluators collect information about users' experiences with a prototype, system, application or design artifact - and they do it in order to improve design.
LEARNING OUTCOMES
- Explain the key concepts and terms used in evaluation.
- Describe the range of different types of evaluation method.
- Explain how methods suit different purposes, stages and contexts of use.
- Show how evaluators mix and modify methods for novel systems.
- Discuss the practical challenges of doing evaluation.
Integral to the design process
Evaluators collect information about users' or potential users' experiences when interacting with a prototype, a computer system, an application or a design artifact. They do this because it improves design.
- Evaluation focuses on the usability of the system - how easy it is to learn and use.
- And on the user experience when interacting with it - how satisfying, enjoyable or motivating the interaction is.
"User experience (UX) is an approach to product development that incorporates direct user feedback throughout the development cycle (human-centered design) in order to reduce costs and create products and tools that meet user needs and have a high level of usability." - User Experience Professionals Association. Note the cost argument: UX work is justified economically, not only ethically.
Why, what, where and when to evaluate
Iterative design and evaluation is a continuous process, and these four questions frame every study.
| Question | Answer |
|---|---|
| Why | To check users' requirements, that users can use the product, and that they like it. |
| What | A conceptual model, early prototypes of a new system, and later, more complete prototypes. |
| Where | In natural and laboratory settings. |
| When | Throughout design. Finished products can also be evaluated, to collect information that informs new products. |
The four answers all resist narrowing: why is three things not one, what starts before there is a system, where is both settings, and when includes after shipping. Evaluation is bigger than 'test the finished product'.
Controlled, natural, and without users
Every evaluation method the course covers falls into one of exactly three types.
1. Controlled settings that directly involve users
Usability labs and research labs. Conditions are controlled as much as possible, and the same conditions apply to every participant.
2. Natural settings involving users
Online communities and products used in public places. There is often little or no control over what users do, especially in in-the-wild settings.
3. Any setting that does not directly involve users
Consultants and researchers critique the prototypes, and may predict and model how successful they will be when used by users.
Living labs extend the idea of a lab: people's use of technology in their everyday lives can be evaluated there, when such evaluations are too difficult to do in a usability lab. The Aware Home (Abowd et al., 2000) was embedded with a complex network of sensors and audio/video recording devices. More recent examples include whole blocks and cities housing hundreds of people (Verma et al., 2017, in Switzerland). Many citizen science projects such as iNaturalist.com can also be thought of as living labs. The concept of a lab is changing to include other spaces where technology use can be studied in realistic environments.
| Case study type | Example from the deck |
|---|---|
| Classic experiment | An investigation into the physiological responses of players of a computer game. |
| Ethnographic study | Visitors at the Royal Highland Show, directed and tracked using a cell phone app. |
| Crowdsourcing | The opinions and reactions of volunteers - the crowd - inform technology evaluation. |
Usability testing and field studies complement each other. Lab work gives control and comparable measures; field work gives ecological validity and surprises. Neither substitutes for the other.
Writing task instructions
Task instructions must be goal-oriented and non-leading. This is the part of test design that most often invalidates a study before it starts.
| Example | |
|---|---|
| Non-leading (correct) | "Try to buy plane tickets for your family vacation." |
| Leading (wrong) | "Press the red button and then click the first top-right button with an airplane icon..." |
- Task-oriented: "Get a ticket to London on a certain date."
- The sample instructions from the deck: What impression does this website give you compared to other airline websites? / You want to fly to Chicago this January for 5 days - find suitable flight times (return tickets) for you and your family. / Decide on suitable flight times and try to buy the tickets. Tell us if anything will affect you from completing the purchase. / How much luggage can you bring with you on this flight?
A leading instruction tells the user where to look, which is exactly the knowledge the test is supposed to measure. If you name the button, you have tested nothing but their ability to follow directions.
Tester, moderator, observer
Three roles, three sets of rules. The exam asks which rule belongs to which role.
Tester (the participant)
1. Think aloud. 2. Be more than 100% honest about likes and dislikes, and about what is difficult or easy. 3. Remember it is not an IQ test. 4. Give suggestions.
Moderator
1. Act as a tour guide. 2. Interrupt only when necessary. 3. Don't answer questions for the tester. 4. Ask questions starting with "What" or "How" if the tester becomes quiet.
Observer
1. Watch. 2. Listen. 3. Take notes.
Testing can be done in person, with the three roles in one of two room setups, or remotely by crowdsourcing - conducting the test with crowdsourced testers so clients can watch and listen. In the remote case the people involved are the observer and the tester.
Participants need to be told why the evaluation is being done, what they will be asked to do, and their rights. Informed consent forms provide this information and act as a contract between participants and researchers. The design of the form, the evaluation process, the data analysis and the data storage methods are typically approved by a higher authority such as an Institutional Review Board.
The moderator's job is defined entirely by restraint: guide, don't interrupt, don't answer, and when the silence gets awkward ask what or how - never why don't you click there.
From session to finding
The lecture gives a four-step common practice, and five things to watch when interpreting data.
- Transcribe the evaluation sessionsVerbal and actions, from both the tester and the moderator.
- Analyze the transcriptBased on the objective of the evaluation - usability, user behaviour and so on.
- Highlight other issuesAnything else observed during the sessions.
- Report findingsPresent what the data supports.
| Consideration | The question it asks |
|---|---|
| Reliability | Does the method produce the same results on separate occasions? |
| Validity | Does the method measure what it is intended to measure? |
| Ecological validity | Does the environment of the evaluation distort the results? |
| Biases | Are there biases that distort the results? |
| Scope | How generalizable are the results? |
Plain validity asks whether the method measures the right thing. Ecological validity asks whether the setting distorts the result - a lab task done perfectly in silence may fail on a noisy train. They are separate marks.
The summary the deck itself gives
Both summary slides in this deck are quotable, and both are examinable.
- Evaluation and design are closely integrated in user-centered design.
- Some of the same techniques are used in evaluation as for establishing requirements, but they are used differently - observation, interviews and questionnaires.
- Three types of evaluation: laboratory based with users, in the field with users, and studies that do not involve users.
- The main methods are observing, asking users, asking experts, user testing, inspection, modeling users' task performance, and analytics.
- Dealing with constraints is an important skill for evaluators to develop.
The six closing rules
Test instructions: goal oriented. Usability testing can be done easily. User's action speaks louder than words. Testing can be done in any stage. You are not your user. Testing with one user is better than none.
"You are not your user" is the entire course compressed into five words - it is Lecture 1's frame-of-reference problem stated as an evaluation rule. Pair it with "action speaks louder than words": what users do in a test outranks what they say in a questionnaire.
Mistakes students usually make
Each claim below is the wrong answer; the line beneath it is the correction, in the wording this course marks against.
Shortest correct answers
The night-before table: every term in this lecture with the smallest answer that still earns the mark.
| Concept | Shortest correct answer |
|---|---|
| Evaluation | Collecting information about users' experiences with a prototype, system, application or artifact, in order to improve design. |
| Why evaluate | To check users' requirements, that users can use the product, and that they like it. |
| What to evaluate | A conceptual model, early prototypes, and later more complete prototypes. |
| Where | Natural and laboratory settings. |
| When | Throughout design; finished products too, to inform new products. |
| Three types | Controlled settings with users; natural settings with users; any setting without users. |
| Living lab | Evaluating everyday technology use in realistic environments - the Aware Home, city blocks, citizen science. |
| Task instructions | Goal oriented and non-leading. |
| Tester rules | Think aloud, be honest, remember it is not an IQ test, give suggestions. |
| Moderator rules | Tour guide, interrupt only when necessary, don't answer questions, ask What or How when the tester goes quiet. |
| Observer rules | Watch, listen, take notes. |
| Informed consent | Tells participants why, what they will do, and their rights; acts as a contract; typically approved by an IRB. |
| Reliability | Does the method produce the same results on separate occasions? |
| Validity | Does the method measure what it is intended to measure? |
| Ecological validity | Does the environment of the evaluation distort the results? |
| Main methods | Observing, asking users, asking experts, user testing, inspection, modeling task performance, analytics. |
Exam-style application
Write your own answer first, then open the model answer. These are the longer-form questions this material generates.
Design a usability test session for a new pharmacy app. Specify the type of evaluation, three task instructions and the roles.
A colleague reports that 8 of 10 users rated the interface 4/5, so the redesign is a success. Critique this using the interpretation considerations.
Explain why the same techniques appear in both the requirements and evaluation lectures, and what changes between them.
Check yourself
6 questions. Every option is explained after submitting, including why the wrong ones are wrong.