SE365 · LECTURE 11Evaluation (Part 1)
SE365 · LECTURE 11 · SLIDE BREAKDOWN

You are not your user.
Testing with one is better than none.

Evaluation is integral to the design process. Evaluators collect information about users' experiences with a prototype, system, application or design artifact - and they do it in order to improve design.

Why/what/where/when3 types of evaluationLiving labsTask instructionsTester / moderator / observerReliability & validity

LEARNING OUTCOMES

  1. Explain the key concepts and terms used in evaluation.
  2. Describe the range of different types of evaluation method.
  3. Explain how methods suit different purposes, stages and contexts of use.
  4. Show how evaluators mix and modify methods for novel systems.
  5. Discuss the practical challenges of doing evaluation.
01 - WHAT EVALUATION IS

Integral to the design process

Evaluators collect information about users' or potential users' experiences when interacting with a prototype, a computer system, an application or a design artifact. They do this because it improves design.

THE UXPA DEFINITION

"User experience (UX) is an approach to product development that incorporates direct user feedback throughout the development cycle (human-centered design) in order to reduce costs and create products and tools that meet user needs and have a high level of usability." - User Experience Professionals Association. Note the cost argument: UX work is justified economically, not only ethically.

02 - THE FOUR QUESTIONS

Why, what, where and when to evaluate

Iterative design and evaluation is a continuous process, and these four questions frame every study.

QuestionAnswer
WhyTo check users' requirements, that users can use the product, and that they like it.
WhatA conceptual model, early prototypes of a new system, and later, more complete prototypes.
WhereIn natural and laboratory settings.
WhenThroughout design. Finished products can also be evaluated, to collect information that informs new products.
MEMORY HOOK

The four answers all resist narrowing: why is three things not one, what starts before there is a system, where is both settings, and when includes after shipping. Evaluation is bigger than 'test the finished product'.

03 - THE THREE TYPES

Controlled, natural, and without users

Every evaluation method the course covers falls into one of exactly three types.

1. Controlled settings that directly involve users

Usability labs and research labs. Conditions are controlled as much as possible, and the same conditions apply to every participant.

2. Natural settings involving users

Online communities and products used in public places. There is often little or no control over what users do, especially in in-the-wild settings.

3. Any setting that does not directly involve users

Consultants and researchers critique the prototypes, and may predict and model how successful they will be when used by users.

Living labs extend the idea of a lab: people's use of technology in their everyday lives can be evaluated there, when such evaluations are too difficult to do in a usability lab. The Aware Home (Abowd et al., 2000) was embedded with a complex network of sensors and audio/video recording devices. More recent examples include whole blocks and cities housing hundreds of people (Verma et al., 2017, in Switzerland). Many citizen science projects such as iNaturalist.com can also be thought of as living labs. The concept of a lab is changing to include other spaces where technology use can be studied in realistic environments.

Case study typeExample from the deck
Classic experimentAn investigation into the physiological responses of players of a computer game.
Ethnographic studyVisitors at the Royal Highland Show, directed and tracked using a cell phone app.
CrowdsourcingThe opinions and reactions of volunteers - the crowd - inform technology evaluation.
COMPLEMENTARY, NOT COMPETING

Usability testing and field studies complement each other. Lab work gives control and comparable measures; field work gives ecological validity and surprises. Neither substitutes for the other.

04 - CASE STUDY: AIRLINE BOOKING

Writing task instructions

Task instructions must be goal-oriented and non-leading. This is the part of test design that most often invalidates a study before it starts.

Example
Non-leading (correct)"Try to buy plane tickets for your family vacation."
Leading (wrong)"Press the red button and then click the first top-right button with an airplane icon..."
WHY LEADING TASKS RUIN A TEST

A leading instruction tells the user where to look, which is exactly the knowledge the test is supposed to measure. If you name the button, you have tested nothing but their ability to follow directions.

05 - RUNNING THE SESSION

Tester, moderator, observer

Three roles, three sets of rules. The exam asks which rule belongs to which role.

Tester (the participant)

1. Think aloud. 2. Be more than 100% honest about likes and dislikes, and about what is difficult or easy. 3. Remember it is not an IQ test. 4. Give suggestions.

Moderator

1. Act as a tour guide. 2. Interrupt only when necessary. 3. Don't answer questions for the tester. 4. Ask questions starting with "What" or "How" if the tester becomes quiet.

Observer

1. Watch. 2. Listen. 3. Take notes.

Testing can be done in person, with the three roles in one of two room setups, or remotely by crowdsourcing - conducting the test with crowdsourced testers so clients can watch and listen. In the remote case the people involved are the observer and the tester.

PARTICIPANTS' RIGHTS AND CONSENT

Participants need to be told why the evaluation is being done, what they will be asked to do, and their rights. Informed consent forms provide this information and act as a contract between participants and researchers. The design of the form, the evaluation process, the data analysis and the data storage methods are typically approved by a higher authority such as an Institutional Review Board.

MEMORY HOOK

The moderator's job is defined entirely by restraint: guide, don't interrupt, don't answer, and when the silence gets awkward ask what or how - never why don't you click there.

06 - ANALYSIS AND INTERPRETATION

From session to finding

The lecture gives a four-step common practice, and five things to watch when interpreting data.

  1. Transcribe the evaluation sessionsVerbal and actions, from both the tester and the moderator.
  2. Analyze the transcriptBased on the objective of the evaluation - usability, user behaviour and so on.
  3. Highlight other issuesAnything else observed during the sessions.
  4. Report findingsPresent what the data supports.
ConsiderationThe question it asks
ReliabilityDoes the method produce the same results on separate occasions?
ValidityDoes the method measure what it is intended to measure?
Ecological validityDoes the environment of the evaluation distort the results?
BiasesAre there biases that distort the results?
ScopeHow generalizable are the results?
VALIDITY vs ECOLOGICAL VALIDITY

Plain validity asks whether the method measures the right thing. Ecological validity asks whether the setting distorts the result - a lab task done perfectly in silence may fail on a noisy train. They are separate marks.

07 - KEY POINTS

The summary the deck itself gives

Both summary slides in this deck are quotable, and both are examinable.

The six closing rules

Test instructions: goal oriented. Usability testing can be done easily. User's action speaks louder than words. Testing can be done in any stage. You are not your user. Testing with one user is better than none.

MEMORY HOOK

"You are not your user" is the entire course compressed into five words - it is Lecture 1's frame-of-reference problem stated as an evaluation rule. Pair it with "action speaks louder than words": what users do in a test outranks what they say in a questionnaire.

MISTAKES STUDENTS USUALLY MAKE

Mistakes students usually make

Each claim below is the wrong answer; the line beneath it is the correction, in the wording this course marks against.

Naming only two types of evaluation.
There are three: controlled settings that directly involve users; natural settings involving users; and any setting that does not directly involve users.
Confusing validity with ecological validity.
Validity asks whether the method measures what it is intended to measure. Ecological validity asks whether the environment of the evaluation distorts the results.
Confusing reliability with validity.
Reliability asks whether the method produces the same results on separate occasions. A method can be perfectly reliable and still measure the wrong thing.
Writing leading task instructions.
"Press the red button then click the airplane icon" tests nothing. The non-leading version is "try to buy plane tickets for your family vacation".
Letting the moderator answer the tester's questions.
The moderator is a tour guide who interrupts only when necessary, does not answer questions for the tester, and asks "What" or "How" questions if the tester goes quiet.
"A living lab is just a bigger usability lab."
Living labs evaluate people's use of technology in their everyday lives, for evaluations too difficult to do in a usability lab - the Aware Home, whole city blocks, and citizen science projects.
Treating consent as a formality.
Informed consent forms tell participants why the evaluation is being done, what they will do and what their rights are, and act as a contract. The form, process, analysis and storage methods are typically approved by a body such as an Institutional Review Board.
CHEAT SHEET

Shortest correct answers

The night-before table: every term in this lecture with the smallest answer that still earns the mark.

ConceptShortest correct answer
EvaluationCollecting information about users' experiences with a prototype, system, application or artifact, in order to improve design.
Why evaluateTo check users' requirements, that users can use the product, and that they like it.
What to evaluateA conceptual model, early prototypes, and later more complete prototypes.
WhereNatural and laboratory settings.
WhenThroughout design; finished products too, to inform new products.
Three typesControlled settings with users; natural settings with users; any setting without users.
Living labEvaluating everyday technology use in realistic environments - the Aware Home, city blocks, citizen science.
Task instructionsGoal oriented and non-leading.
Tester rulesThink aloud, be honest, remember it is not an IQ test, give suggestions.
Moderator rulesTour guide, interrupt only when necessary, don't answer questions, ask What or How when the tester goes quiet.
Observer rulesWatch, listen, take notes.
Informed consentTells participants why, what they will do, and their rights; acts as a contract; typically approved by an IRB.
ReliabilityDoes the method produce the same results on separate occasions?
ValidityDoes the method measure what it is intended to measure?
Ecological validityDoes the environment of the evaluation distort the results?
Main methodsObserving, asking users, asking experts, user testing, inspection, modeling task performance, analytics.
APPLY IT

Exam-style application

Write your own answer first, then open the model answer. These are the longer-form questions this material generates.

Design a usability test session for a new pharmacy app. Specify the type of evaluation, three task instructions and the roles.
Type: a controlled setting that directly involves users, so conditions are the same for every participant and errors and times are comparable - complemented later by a field study, since the two complement each other. Tasks (goal-oriented, non-leading): (1) "You need to refill a prescription you collected last month - do that." (2) "Find out whether this medicine can be taken with the one you already have." (3) "What impression does this app give you compared with others you have used?" None names a button or a screen. Roles: the tester thinks aloud, is more than honest about likes, dislikes and difficulty, is reminded it is not an IQ test, and is invited to suggest improvements; the moderator acts as a tour guide, interrupts only when necessary, never answers the tester's questions, and asks "What are you thinking?" or "How would you expect that to work?" during silences; the observer watches, listens and takes notes. Before any of it, an informed consent form tells participants why the study is happening, what they will do and what their rights are.
A colleague reports that 8 of 10 users rated the interface 4/5, so the redesign is a success. Critique this using the interpretation considerations.
Validity: a satisfaction rating measures how users feel about the interface, not whether they could use it - if the study's objective was usability, the method does not measure what it was intended to measure. The deck's own rule applies: the user's action speaks louder than words, so completion rates, times and error counts should carry the finding, with the rating as supporting data. Reliability: would the same ratings appear on another occasion, or did the novelty of the redesign inflate them? Ecological validity: if the ratings came from a quiet lab, they may not survive the real setting. Bias: ratings given to the person who built the interface are systematically generous, which is one reason the moderator should not be the designer. Scope: ten participants cannot support a general claim about the user population. Report what the data supports and no more.
Explain why the same techniques appear in both the requirements and evaluation lectures, and what changes between them.
Observation, interviews and questionnaires appear in both, because both activities need evidence about real users - but the deck stresses that they are used differently. In requirements the goal is to understand users, tasks and context in order to produce a stable set of requirements, so the questions are open and exploratory and the setting is the user's own world. In evaluation the goal is to judge a specific artifact - a conceptual model, an early prototype or a more complete one - against agreed usability and UX goals, so tasks are predefined and goal-oriented, conditions are held constant across participants where control is wanted, and performance is measured rather than described. The framing questions also change: evaluation asks why, what, where and when, and the answers span both natural and laboratory settings and continue after the product ships, where findings feed into new products.
POP QUIZ

Check yourself

6 questions. Every option is explained after submitting, including why the wrong ones are wrong.