STAT101 - Introduction to Probability Theory and Statistics
Chapter 1 Slide Breakdown โ Groebner, Shannon & Fry, Business Statistics, 10th Global Edition
Chapter 1 lays the groundwork for the entire course. It answers four questions: (1) what is business statistics and how is it split into descriptive vs. inferential branches, (2) what are the main ways to collect data, (3) what is the difference between a population and a sample and how do we pick a sample fairly, and (4) once we have data, how do we classify it so we know which statistical tools apply to it. Almost every later chapter assumes you are fluent in this vocabulary, so exam questions here are heavy on definitions and "identify which type" scenarios rather than calculations.
Business statistics is a collection of procedures and techniques used to convert data into meaningful information in a business environment.
The flow is always: Data โ Statistical Procedures (Descriptive or Inferential) โ Information.
A data set is organized like a spreadsheet:
| Branch | Purpose | Example |
|---|---|---|
| Descriptive Statistics | Procedures and techniques designed to describe data (summarize what you already have โ no generalizing beyond it). | Charts, graphs, tables; numerical measures like the average. |
| Inferential Statistics | Tools and techniques that help decision makers draw inferences from a sample about a larger population. | Estimation (e.g., estimating average family income from a sample) and Hypothesis Testing (e.g., testing a claim that income exceeds $45,000). |
This is the most basic descriptive numerical measure โ you will meet dozens of variations on it (weighted mean, CV, percentiles) in later chapters.
Scenario A: A store manager creates a bar chart of monthly sales by department. โ This is descriptive statistics โ it only summarizes the data already collected, nothing is being generalized.
Scenario B: A pollster surveys 500 out of 2 million registered voters and uses the sample result to estimate what percentage of ALL voters support a candidate. โ This is inferential statistics โ a sample statistic is being used to infer something about the whole population.
Ask yourself: "Is the conclusion limited strictly to the data in front of me, or am I making a claim about something bigger than what I measured?" If it stays within the data โ descriptive. If it reaches beyond the sample to the population โ inferential. The two flavors of inferential statistics are estimation ("what is the value?") and hypothesis testing ("is this claim true?").
Students often think "using a formula" automatically means inferential statistics. Not true โ computing an average, a percentage, or drawing a chart is still descriptive, even though it uses a formula. What makes something inferential is generalizing from a sample to a population.
Data Collection Techniques branch into four categories: Experiments, Telephone Surveys, Written Questionnaires and Surveys, and Direct Observation and Personal Interview.
Experiment: a process that produces a single outcome whose result cannot be predicted with certainty.
Experimental design: a plan for performing an experiment in which the variable of interest is defined. It specifies ahead of time what will be measured and how conditions will be controlled.
A food-science team wants to know how blanching time (10, 15, 20, 25 minutes), blanching temperature (100ยฐ, 110ยฐ, 120ยฐ), and potato category (1โ4) affect the quality of French fries. They run controlled trials for every combination and measure the outcome (e.g., crispness score). This is an experiment because the researchers actively control the input conditions rather than just observing what happens naturally โ that is what "experimental design" means in practice.
Political polls, market research, customer satisfaction studies, healthcare studies.
Respondents screen calls / won't answer, respondents not home, cell numbers must be manually dialed, surveys must stay short or respondents hang up, requires trained callers.
Define the Issue โ Define the Population of Interest โ Develop Survey Questions โ Pretest the Survey โ Determine Sample Size and Sampling Method โ Select Sample and Make Calls.
| Question Type | Definition | Example |
|---|---|---|
| Closed-End Questions | Require the respondent to select from a short list of defined choices. | "What is your preferred rental car company?" ______ (pick from list) |
| Demographic Questions | Questions relating to the respondents' characteristics, backgrounds, and attributes. | "What is your gender?" / "List your age on your last birthday?" |
Must be carefully formatted; need a motivational statement of purpose; can more easily allow for open-ended questions; less expensive to administer.
Define the Issue โ Define the Population of Interest โ Design the Survey Instrument โ Pretest the Survey โ Determine Sample Size and Sampling Method โ Select Sample and Send Surveys.
"Under what circumstances would you favor a reduction in the Federal long-term capital gains tax rate?" ______________ โ respondents answer in their own words, with no predefined list of choices (contrast this with a closed-end question).
| Advantages | Disadvantages | |
|---|---|---|
| Written Questionnaires | Questions are standardized for all respondents; if online, data are captured electronically; potentially allows for more questions. | Responses may not be timely; potentially low response rate (5โ10%); no chance to clarify questions. |
Typical applications: employee surveys, college course/professor evaluations, customer satisfaction studies, academic research.
Data are physically observed and recorded based on what takes place in the process. Subjective and time-consuming.
Applications: customer product preferences, seatbelt-compliance rates, company safety-policy compliance.
Challenges: avoiding the "Hawthorne Effect" (people behave differently when they know they're being watched), maintaining consistency among observers, needing properly trained and unbiased observers; costly and time-consuming.
Structured: verbally administered, a specific set of predetermined questions asked in the same order to all respondents, no follow-up or elaboration.
Unstructured: no predetermined set of questions; begins with a general question and lets discussion go in any direction; interviewer probes for clarification.
Applications: customer satisfaction, supplier qualification, health-status measurement.
Challenges: interviewer bias, capturing responses accurately, time requirements, limited number of interviewees, interviewer training and consistency.
| Data Collection Method | Advantages | Disadvantages |
|---|---|---|
| Experiments | Provide controls; preplanned objectives | Costly; time-consuming; requires planning |
| Telephone Surveys | Timely; relatively inexpensive | Poor reputation; limited scope and length |
| Mail Questionnaires / Written Surveys | Inexpensive; can expand length; can use open-end questions | Low response rate; requires exceptional clarity |
| Direct Observation / Personal Interview | Expands analysis opportunities; no respondent bias | Potential observer bias; costly |
MCQs love to swap an advantage from one row with another row (e.g., saying "telephone surveys have no respondent bias"). Learn which advantage/disadvantage is uniquely tied to which method: cost and control โ experiments; speed and cheapness (but "poor reputation") โ phone; low response rate โ written; observer bias โ observation/interview.
A group interview using a facilitator. Obtains data about the group's perspectives and opinions; responses are coded into categories for analysis; members share some common attribute.
Studying existing (secondary) data contained in databases, financial records, annual reports, etc. Potentially less expensive than generating primary data, but may not be complete.
A bar code provides information about a product (price, supplier, etc.); a bar code scanner automatically captures the data.
Radio Frequency Identification Device โ transmitting devices attached to a product; an RFID receiver reads the data signal. Holds more data than a bar code, does not require line of sight, and works better than bar codes in harsh environments โ but signals can be compromised.
When collecting data (from any method above) you must watch for: Data Accuracy, Interviewer Bias, Nonresponsive Bias (people who don't respond may differ systematically from those who do), Selection Bias (the way units are chosen skews results), Observer Bias, and Measurement Error. These feed into two big-picture concerns: Internal Validity (did the study actually measure what it claims to measure, free of confounding effects?) and External Validity (can the results be generalized beyond this specific study?).
| Term | Definition |
|---|---|
| Population | The set of all objects or individuals of interest, or the measurements obtained from all objects or individuals of interest. |
| Sample | A subset of the population. |
| Census | An enumeration of the entire set of measurements taken from the whole population. |
If the population is "All Customers in the Market Area" (say, letters a through z), the sample is just a subset of those customers (say, b, c, g, h, k, l, m, n, o, r, s, v, w, z) โ chosen using one of the techniques below.
| Population (Parameter) | Sample (Statistic) | |
|---|---|---|
| Mean | ฮผ (mu) | xฬ (x-bar) |
| Proportion | p | pฬ (p-bar) |
Parameter: a descriptive numerical measure (like an average or proportion) computed from an entire population. Example: the average yards gained per play by all NFL teams in a season.
Statistic: a descriptive numerical measure computed from a sample selected from a population. Example: the average credits taken by a sample of students at a university.
Sampling error is the natural, unavoidable difference between a sample statistic (like xฬ) and the true population parameter (like ฮผ) that occurs simply because you measured a subset instead of the whole population โ even a perfectly executed random sample will have some sampling error. Non-sampling error is everything else that can make data wrong: interviewer bias, measurement error, nonresponse bias, data entry mistakes, poorly worded questions. Non-sampling error can be reduced or eliminated with better procedures; sampling error can only be reduced (never eliminated) by increasing the sample size.
| Type | Definition |
|---|---|
| Statistical (Probability) Sampling | Sampling methods that use selection techniques based on chance selection โ every item in the population has a known or calculable chance of being included. |
| Nonstatistical Sampling | Methods of selecting samples that use convenience, judgment, or other non-chance processes. |
Collected in whatever manner is most convenient for the researcher.
Example: a university gathers opinions on switching from a quarter to a semester system by stopping students who happen to enter the library.
Based on judgments about who in the population is most likely to provide the needed information.
Example: a manager selects her ten largest suppliers to interview about satisfaction with the ordering process.
The sample size selected from a given segment is proportional to the number of items in the population belonging to that segment.
Example: if the target population is 60% female and 30% college graduates, the sample is deliberately built to also be 60% female and 30% college graduates.
Every possible sample of a given size has an equal chance of being selected. Selection may be with or without replacement. Obtained via a random number table or generator (e.g., Excel's Data Analysis โ Random Number Generation).
Divide the population into subgroups (strata) by a common characteristic (e.g., gender, income level), then select a simple random sample from each stratum and combine them into one sample.
Decide sample size n; divide the ordered population of N individuals into groups of k individuals where k = N/n; randomly pick one individual from the first group, then take every k-th individual after that.
Divide the population into several "clusters," each representative of the whole population (e.g., county, warehouse bin). Randomly select a sample of clusters, then use every item in the chosen clusters (or sample within them).
where N = population size, n = desired sample size, k = the sampling interval ("select every k-th item").
The population is "Cash Holdings of All Financial Institutions in the U.S." It is stratified into Stratum 1 = Large Institutions, Stratum 2 = Medium-Size Institutions, Stratum 3 = Small Institutions. A simple random sample is drawn separately from each stratum (nโ, nโ, nโ), and the three sub-samples are combined. Advantage: if the strata are correctly determined, data within each stratum will be fairly homogeneous, so the total sample size needed (and hence sampling cost) is smaller than an equivalent simple random sample.
A population has N = 64 people and we want a sample of n = 8. Then k = N/n = 64/8 = 8. We randomly pick one person from the first group of 8 (say, person #3), then take every 8th person after that: #3, #11, #19, #27, #35, #43, #51, #59.
A warehouse manager wants a statistical sample of products spread across a warehouse in 10 aisles ร 3 stack heights ร 20 stacks per aisle = 600 bins. Each bin is a cluster. She randomly selects 5 clusters (bins) and samples every item in those 5 bins โ much cheaper than randomly sampling individual items scattered across all 600 bins.
Ask: "Did they divide into groups by a shared trait and sample from EVERY group?" โ Stratified. "Did they pick every k-th item off an ordered list?" โ Systematic. "Did they randomly pick whole groups and use everyone/everything inside them?" โ Cluster. "Is every individual item equally likely, with no grouping at all?" โ Simple Random.
The starting point in analyzing data is to know exactly what kind of data you have collected โ this determines which charts, summary measures, and statistical tests are valid to use on it.
| Type | Definition | Subtypes / Examples |
|---|---|---|
| Quantitative | Measurements whose values are inherently numerical. | Discrete: countable, gaps between possible values, no values available in between (e.g., number of children). Continuous: infinitely many possible values, no gaps (e.g., weight, volume). |
| Qualitative | Data whose measurement scale is inherently categorical. | Marital status, political affiliation, eye color. |
| Type | Definition | Example |
|---|---|---|
| Time-Series | A set of consecutive data values observed at successive points in time. | Stock price recorded daily for a year; Atlanta's sales across 2009โ2012. |
| Cross-Sectional | A set of data values observed at a fixed point in time. | Bank data about all its loan customers today; all four cities' sales in 2009 only. |
In a table with cities down the rows and years across the columns: reading across one row (same city, changing years) gives you time-series data. Reading down one column (same year, changing cities) gives you cross-sectional data.
Measurement levels form a hierarchy from lowest to highest information content: Nominal โ Ordinal โ Interval โ Ratio. Each higher level can do everything the lower levels can do, plus more.
Categorical codes, ID numbers, category names โ labels only, no meaningful order or math.
Examples: Student ID number, favorite color, gender, marital status.
Rankings / ordered categories โ order matters, but the gaps between values aren't necessarily equal or meaningful.
Examples: College class standing (UG=1, Grad=2); Age group (Under 18=1, 18โ30=2, 31โ60=3, Over 60=4); Satisfaction level (VS=1, S=2, Neutral=3, D=4, VD=5).
Numerical, ordered, equal gaps between values โ but no true zero (zero doesn't mean "none").
Example: Temperature in ยฐC or ยฐF (22ยฐ, 87ยฐ, 0ยฐ, 45.92ยฐ) โ 0ยฐ does not mean "no temperature," so ratios don't make sense (20ยฐ is not "twice as hot" as 10ยฐ).
Numerical, ordered, equal gaps, AND a true zero point โ full mathematical operations (including ratios) are meaningful.
Examples: Weight (1.35 oz.), Time (33.05 sec.), Pay rate per hour ($33.50), Interest rates (4.05%).
To tell Interval from Ratio, ask: "Does a value of zero mean the complete absence of the quantity?" If yes โ Ratio (0 weight = no weight, so you CAN say "10 lbs is twice 5 lbs"). If no โ Interval (0ยฐC โ no temperature, so you CANNOT say "20ยฐC is twice as hot as 10ยฐC"). To tell Ordinal from Nominal, ask: "Is there a meaningful order/ranking?" If yes โ Ordinal. If the numbers are just labels (like a Student ID) โ Nominal, even though it "looks numeric."
Don't assume that if data is stored as a number, it must be interval or ratio. A student ID number, a jersey number, or a ZIP code are all numeric-looking but are actually nominal โ the numbers are just labels with no arithmetic meaning (adding two student IDs is meaningless).
A spreadsheet of colleges has columns: College Name, State, Public(1)/Private(2), Math SAT, Verbal SAT, # applications received, # applications accepted, # new students enrolled, # full-time undergrad, # part-time undergrad.
1. Identify each factor (variable) in the data set. 2. Determine whether the data are time-series or cross-sectional. 3. Determine which factors are quantitative vs. qualitative. 4. Determine the level of data measurement (nominal, ordinal, interval, ratio) for each factor.
1. A company computes the average satisfaction score from the customer survey it collected. This is an example of:
2. Which data collection method is most likely to suffer from the "Hawthorne Effect"?
3. A researcher divides a population into strata by income level and takes a simple random sample from each stratum. This is:
4. A population has N = 500 and a sample of n = 50 is needed using systematic sampling. What is the sampling interval k?
5. Temperature measured in degrees Celsius is an example of which data measurement level?
6. Which of the following is a nonstatistical (nonrandom) sampling technique?
7. A student's ID number is best classified as:
8. Explain the difference between sampling error and non-sampling error, and state which one can be reduced by increasing sample size.
Answer: Sampling error is the unavoidable difference between a sample statistic and the true population parameter that arises simply because a subset (not the whole population) was measured โ it exists even with a perfectly executed random sample. Non-sampling error covers all other sources of error: interviewer bias, measurement error, nonresponse bias, poorly worded questions, or data entry mistakes. Increasing the sample size reduces sampling error (though never eliminates it); non-sampling error must be reduced through better procedures and design, not by adding more observations.
9. A company wants to know the average number of hours 10,000 employees spend on training each year. It surveys 200 randomly chosen employees and finds their average is 14.5 hours. Identify the population, the sample, the parameter of interest, and the statistic being computed.
Answer: Population = all 10,000 employees. Sample = the 200 employees surveyed. Parameter of interest = ฮผ, the true (unknown) average training hours for all 10,000 employees. Statistic = xฬ = 14.5 hours, the average computed from the sample, used as a point estimate of ฮผ.
10. A university wants to survey students but only stops students who happen to be walking by the student center between 12โ1pm. Which nonstatistical sampling method is this, and what is the main risk with this approach?
Answer: This is convenience sampling โ data collected in whatever manner is easiest for the researcher. The main risk is that the sample may not represent the full population well (e.g., only students with lunch-time classes near the student center are captured), introducing selection bias into the results.
where N = population size, n = desired sample size.
| Term | Definition |
|---|---|
| Business Statistics | Procedures and techniques used to convert data into meaningful information in a business environment. |
| Descriptive Statistics | Procedures and techniques designed to describe data. |
| Inferential Statistics | Tools and techniques that help decision makers draw inferences from a sample about a population. |
| Estimation | An inferential procedure used to estimate an unknown population value from sample data. |
| Hypothesis Testing | An inferential procedure that uses sample evidence to test a claim about a population. |
| Experiment | A process that produces a single outcome whose result cannot be predicted with certainty. |
| Experimental Design | A plan for performing an experiment in which the variable of interest is defined. |
| Closed-End Question | A question requiring the respondent to choose from a short list of defined choices. |
| Open-End Question | A question allowing respondents to answer freely in their own words. |
| Demographic Question | A question about the respondent's characteristics, background, or attributes. |
| Structured Interview | A verbally administered interview using a predetermined, fixed-order set of questions with no follow-up. |
| Unstructured Interview | An interview that begins with a general question and allows discussion to go in any direction, with probing follow-ups. |
| Focus Group | A facilitated group interview used to obtain shared perspectives and opinions. |
| Hawthorne Effect | The tendency for people to change their behavior because they know they are being observed. |
| Population | The set of all objects or individuals of interest, or the measurements obtained from all of them. |
| Sample | A subset of the population. |
| Census | An enumeration of the entire set of measurements from the whole population. |
| Parameter | A descriptive numerical measure computed from an entire population (e.g., ฮผ, p). |
| Statistic | A descriptive numerical measure computed from a sample (e.g., xฬ, pฬ). |
| Sampling Error | The unavoidable difference between a sample statistic and the true population parameter, caused solely by sampling a subset instead of the whole population. |
| Non-Sampling Error | Error caused by factors other than sampling itself, such as interviewer bias, measurement error, or nonresponse โ reducible through better procedures. |
| Statistical (Probability) Sampling | Sampling based on chance selection, where every item has a known or calculable chance of inclusion. |
| Nonstatistical Sampling | Sampling based on convenience, judgment, or other non-chance processes. |
| Convenience Sampling | Nonstatistical sampling collected in whatever manner is most convenient for the researcher. |
| Judgment Sampling | Nonstatistical sampling based on judgment about who is most likely to provide needed information. |
| Ratio Sampling | Nonstatistical sampling where the sample proportion from each segment matches that segment's proportion in the population. |
| Simple Random Sampling | Statistical sampling where every possible sample of a given size has an equal chance of selection. |
| Stratified Random Sampling | Dividing the population into strata by a shared characteristic and randomly sampling from each stratum. |
| Systematic Random Sampling | Selecting every k-th individual from an ordered population after a random start, where k = N/n. |
| Cluster Sampling | Dividing the population into representative clusters, randomly selecting clusters, and sampling all (or some) items within them. |
| Quantitative Data | Data whose values are inherently numerical (discrete or continuous). |
| Qualitative Data | Data whose measurement scale is inherently categorical. |
| Discrete Data | Countable quantitative data with gaps between possible values. |
| Continuous Data | Quantitative data that can take infinitely many values with no gaps. |
| Time-Series Data | Consecutive data values observed at successive points in time. |
| Cross-Sectional Data | Data values observed at a single, fixed point in time. |
| Nominal Data | The lowest measurement level: categorical codes/labels with no inherent order. |
| Ordinal Data | Data with a meaningful rank order, but unequal or undefined gaps between values. |
| Interval Data | Numerical data with equal gaps between values but no true zero point. |
| Ratio Data | The highest measurement level: numerical data with equal gaps and a true zero point. |