1 of 60

2 of 60

3 of 60

Please read this disclaimer before proceeding:

This document is confidential and intended solely for the educational purpose of RMK Group of Educational Institutions. If you have received this document through email in error, please notify the system manager. This document contains proprietary information and is intended only to the respective group / learning community as intended. If you are not the addressee you should not disseminate, distribute or copy through e-mail. Please notify the sender immediately by e-mail if you have received this document by mistake and delete this document from your system. If you are not the intended recipient you are notified that disclosing, copying, distributing or taking any action in reliance on the contents of this information is strictly prohibited.

4 of 60

MARKETING RESEARCH AND MARKETING�MANAGEMENT�(20CB406)

Department: CSBS�Batch/Year: II YEAR / IV SEM�Created by:Dr.S.D. Uma Mageswari�Date: 07.03.2022�

5 of 60

UNIT IV

  • Marketing Research: Introduction, Type of Market Research, Scope, Objectives & Limitations Marketing Research Techniques, Survey Questionnaire design & drafting, Pricing Research, Media Research, Qualitative Research
  • Data Analysis: Use of various statistical tools – Descriptive & Inference Statistics, Statistical Hypothesis Testing, Multivariate Analysis - Discriminant Analysis, Cluster Analysis, Segmenting and Positioning, Factor Analysis

6 of 60

UNIT IV

MARKETING RESEARCH

9

Marketing Research: Introduction, Type of Market Research, Scope, Objectives & Limitations Marketing Research Techniques, Survey Questionnaire design & drafting,Pricing Research, Media Research, Qualitative Research

Data Analysis: Use of various statistical tools - Descriptive & Inference Statistics, Statistical Hypothesis Testing, Multivariate Analysis - Discriminant Analysis, Cluster Analysis, Segmenting and Positioning, Factor Analysis

4.1 Market Research and Marketing Research

Generally Market Research and Marketing Research are confused to be the same. But there is a clear distinction between the both.

Market Research: Market Research involves researching a specific industry or market. Ex: Researching the automobile industry to discover the number of competitors and their market share.

Marketing Research: Marketing Research analyses a given marketing opportunity or problem, defines the research and data collection methods required to deal with the problem or take advantage of the opportunity, through to the implementation of the project. It is a more systematic method which aims to discover the root cause for a specific problem within an organisation and put forward solutions to that problem. Ex: Research carried out to analyze and find solution for increasing turnover in an organisation

4.2 Types of Market Research:

7 of 60

4.3 Marketing Research:

□ "The careful and objective study of product design, markets and such transfer activities as physical distribution, warehousing advertising and sales management. Thus the scope of marketing research lies in its variety of applications."

It is a technique to know:

1. Who are customers of our products or services?

2. Where do they live?

3. When and how do they buy the product and services?

4. Are customers of our products satisfied with the products?

5. Who are our main competitors in the market?

6. Are the company's products inferior or superior to competitors' products?

7. What policies and strategies are they following?

4.4 Scope of Marketing Research:

□ The scope of marketing research stretches from the identification of

consumer wants and needs to the evaluation of consumer satisfaction.

SCOPE OF MARKETING RESEARCH

1. Size of the present and potential market.

6. Analysis of market demand.

2. Consumer needs wants, habits and

7. Knowledge of competitors and their

behaviour.

products.

3. Dealer wants and preferences.

8. Knowing the profitability of different

4. Analysis of the market size according to

markets.

age, sex, income, profession, standard of

9. Study the market changes and market

living etc.

conditions.

5. Geographic location of customers.

10. Analysis of various channels of distribution.

8 of 60

4.5 Objectives of Marketing Research:

1. To understand the economic factors affecting the sales volume and their opportunities.

2. To understand the competitive position of rival products.

3. To evaluate the reactions of consumers and customers.

4. To study the price trends.

5. To evaluate the system of distribution.

6. To understand the advantages and limitation of the products.

7. To find new methods of packaging, by comparing other similar packages.

8. To analyze the market size.

9. To know the estimation of demand.

10. To evaluate the profitability of different markets.

11. To study the customer's acceptance of products.

12. To assess the volume of future sales.

13. To study the nature of the market, its location and its potentialities.

14. To find solutions to problems relating to marketing of goods and services.

15. To evaluate policies and plans in the right course of action.

16. To know the development of science and technology.

17. To know the complexity of marketing.

18. To measure the effectiveness of advertising.

19. To estimate the potential market for a new product.

20. To assess the strength and weakness of the competitors.

4.6 Marketing Research Methods

Methodologically, marketing research uses four types of research designs.

• Qualitative marketing research - This is generally used for exploratory purposes. The data collected is qualitative and focuses on people's opinions and attitudes towards a product or service. The respondents are generally few in number and the findings cannot be generalised tot eh whole population. No statistical methods are generally applied.

Ex: Focus groups, In-depth interviews, and Projective techniques

9 of 60

Quantitative marketing research - This is generally used to draw conclusions for a specific problem. It tests a specific hypothesis and uses random sampling techniques so as to infer from the sample to the population. It involves a large number of respondents and analysis is carried out using statistical techniques.

Ex: Surveys and Questionnaires

Observational techniques - The researcher observes social phenomena in their natural setting and draws conclusion from the same. The observations can occur cross-sectionally (observations made at one time) or longitudinally (observations occur over several time-periods)

Ex: Product-use analysis and computer cookie tracing

Experimental techniques - Here, the researcher creates a quasiartificial environment to try to control spurious factors, then manipulates at least one of the variables to get an answer to a research

Ex: Test marketing and Purchase laboratories

4.7 Marketing Research Process

The Marketing Research Process involves a number of inter-related activities which have bearing on each other. Once the need for Marketing Research has been established, broadly it involves the steps as depicted in Figure 1 below:

10 of 60

1. PROBLEM DEFINITION

This is the starting point in the marketing research exercise. In problem definition it is important to be specific, avoiding ambiguities and generalities. Care should also be taken, not to define problems in too narrow a field as that may distract the researcher's perspective. This may even affect creativity in the research.

2. RESEARCH OBJECTIVES

Once the problem is defined, the next logical step is to state what the researcher wants to achieve. This statement is called objectives. To be meaningful and help focus the researcher's attention, these objectives should be specific, attainable & measurable. The purpose of these objectives is to act as a guide to the researcher and help him in maintaining a focus all through the research.

3. RESEARCH DESIGN

The third stage in the marketing research process is deciding on the research design. There are three types of research designs, namely:

1. Exploratory: This kind of research is conducted when the researcher does not know how & why a certain phenomenon occurs. Since the prime goal of an exploratory research is to know the unknown, this research is unstructured. Focus groups, interviewing key customer groups, experts and even search for printed or published information are some common techniques.

11 of 60

2. Descriptive: This research is carried out to describe a phenomenon or market characteristics. For example, a study to understand buyer behavior & describe characteristics of the target market is a descriptive research. Continuing the above example of service quality, a research done on how consumers evaluate the quality of competing service institutions can be considered as an example of descriptive research.

3. Causative: This kind of research is done to establish a cause and effect relationship, for example the influence of income & lifestyle on purchase decision. Here the researcher may like to see the effect of rising income & changing lifestyle on consumption of select products.

4. SOURCES OF DATA

Once the research design has been decided upon, the next stage is that of selecting the sources of data. Essentially there are two sources of data or informationsecondary & primary

• Secondary data: This refers to the information that has been collected earlier by someone else. Often this includes printed or published reports, news items, industry or trade statistics etc. this also includes internal documents like invoices, sales reports, payment history of customers etc. these are important to the researcher as they provide an insight to the problem. Often the preliminary investigation is restricted to secondary data.

• Primary data: To overcome the limitations of incompatibility, obsolescence and bias, the researcher turns to the primary data. This is also resorted to when the secondary data is incomplete. Primary sources refer to data collected directly from the market place- customers, traders & suppliers often are the major sources. They are often reliable data sources and help in overcoming limitations of secondary data. The problem in primary data is its cost, both In terms of money & time, and often a researcher bias also creeps in.

5. DATA COLLECTION

12 of 60

The researcher is now ready to take the plunge. But still he or she needs to be clear about the following.

Procedure for data collection.

Data can be collected through any or combination of the following techniques.

• Observation: This technique involves observing how a customer behaves in the shopping area, how he or she dresses up & what does the customer say when he or she sees the product.

• Experimentation: This is a technique that involves experimenting new product ideas, advertising copies & campaigns, sales promotion ideas & even pricing & distribution strategies with the target customer group. These experiments can be conducted in an uncontrolled environment or in a controlled & simulated market environment.

Tools for data collection

The researcher has to decide on the appropriate tool for data collection. These tools are:-

• Questionnaire — used for the survey method

• Interview schedule — used mainly for exploratory research

• Association test — primarily used in qualitative research, also called as TAT (Thematic Apperception Test)

6. DATA ANALYSIS

The next stage is that of data analysis .It is important to understand raw data has no usage in marketing research .hence appropriate analytical tools must be used. The most elementary is the arithmetic analysis using percentile and ratios. Statistical analysis like mean, median, mode, percentages, standard deviation and coefficient of correlations should be used wherever applicable

7. REPORT & PRESENTATION

13 of 60

The last stage is that of writing out a report and making a presentation to the Decision —maker. It is important that the report has summary, called the executive summary, giving a bird's-eye view of the research. This is because most senior managers have little time for going through the entire report in depth. The executive summary can direct the reader's attention to specific issues by turning to the relevant sections in the report and should not exceed thousand words.

The report should be structured and pages chronologically numbered generally, the structure of a good repot is somewhat like the following:

• Introduction to the problem

• Marketing research finding or survey findings

• Interpretation of research finding

• Policy implications

4.8 PRICING RESEARCH

□ Pricing research is a method of research that measures and evaluates the impact of changes in price of a product on its demand.

□ It is used by organizations to help determine an optimal price for new products, in order to maximise revenue and market share.

□ This type of research is quantitative in nature.

There are two key benefits of conducting pricing research:

(i) the prediction of consumers' response to price changes, and

(ii) the discovery of psychological effects of price points on sales (demand

4.8.1 Methods of pricing research:

a) Van Westendorp Price Sensitivity Meter (PSM)

The Van Westendorp Price Sensitivity Meter constructs a range of acceptable price points for a given product, determining the expected price range at which consumers will be willing to purchase it. This range is constructed by having customers evaluate a product and then respond to the following four questions:

14 of 60

• Too Expensive: "At what price would you begin to think this product is too expensive to consider?"

• Expensive: "At what price would you begin to think this product is expensive but worth considering?"

• Cheap: "At what price would you begin to think this product is a bargain?"

• Too Cheap:"At what price would you begin to think the product is so inexpensive that you would question its quality?"

Once responses are collected, the cumulative frequency of the different answers are charted in order to determine a series of acceptable price points.

These price points will range from a lower threshold to an upper threshold, and will also include the optimal price point.

□ PSM is used to understand customers' pricing expectations, rather than their willingness to pay or their likelihood to buy.

□ It is used to identify how much respondents would expecta product to cost.

b) Gabor-Granger Technique

The Gabor-Granger technique involves testing four to five different price points by asking respondents their likelihood to purchase the product at each one of these points.

Respondents indicate their likelihood to purchase at these predefined price points, and this data is used to determine an optimal price point for the product within the market. The Gabor-Granger technique asks respondents to evaluate predetermined price points that have already been vetted by the company. It identifies the optimal price range for a product, considering it in isolation.

15 of 60

c) Conjoint Analysis

Conjoint Analysis, also known as discrete choice analysis, is a pricing research technique that is considered to be the most reliable way to determine the price of a product.

In this technique, respondents are given a choice of two to five product profiles, each with different configurations. Respondents are asked to choose one of these profiles. The data collected from respondents allows researchers to create pricing and packaging models that are most likely to appeal to customers.

d) Brand-Price Trade-Off (BPTO)

BTPO, or Brand-Price Trade-Off, is a statistical tool that is used to identify the effect of price on different areas such as profitability, revenue, market volume, and brand awareness. It is a choice-based pricing technique that depicts consumers' differing preferences for brands based on their pricing.

Survey respondents are shown a range of branded products, each with a price associated with it. The range usually consists of 3 to 5 products. Consumers are then asked which "offer" would be most appealing to them in a hypothetical buying scenario.

BPTO is useful in situations where you want to understand the relationship between a brand and its prices.

4.9 Media Research

□ Media Research is the study of the effects of the different mass media on social, psychological and physical aspects. □ Research segments the people based on what television programs they watch, radio they listen, media they access and magazines they read.

Parameters in Media Research

1. The nature of medium being used

2. The working of the medium

3. Technologies involved in it

4. Difference and similarities between it and other media vehicles

5. Functions and services provided by it

6. Cost associated and access to new medium

7. Effectiveness and how it can be improved

As decision process depends on data, thus media research has grown to be utilized for long range planning. Research is in growth phase due to competitions between different media.

16 of 60

Importance of Media Research

1) Gives useful information: media research helps to understand and determine new trends and get valuable insights into the field of mass media and communication, which further enables to determine how more people can be reached within a short span of time.

2) Helps frame news better: A thorough media research study helps to understand how news can be framed better and make it more accessible to the target audience. It helps in analysis and composition of views, news, and data.

3) Makes the (Message) story better and more accurate: Thorough media research also helps to create more accurate and objectively apt stories. It is impossible to do so if your efforts are not directed towards investigating each aspect of a story.

Steps involved in an extensive media research study:

  • Pick a problem.
  • Go through currently existing research and theories that are relevant.
  • Come up with well-articulated research questions and a hypothesis or hypotheses.
  • Figure out an apt research design or algorithm and then gather relevant data.
  • Conduct a thorough analysis of the results and determine their feasibility.
  • Present those results in a structured format.
  • Leverage the valuable insights of the study whenever it’s required.

17 of 60

4.9.1. Media Strategy:

□ The usage of the appropriate media mix in order to achieve desired and optimum outcomes from the advertising campaign.

□ It plays a key role in advertising campaigns.

□ Media Strategy is not just about informing customers about products or services but also placing right message towards the right people at the right time.

4.9.2 Importance of Media Strategy

□ 1. Location : Location is all about where to launch and run the campaign. Location should be the one which gives maximum ROI. In current scenarios, online and offline locations are both considered while deciding a media strategy.

□ 2. Budget: For deciding the media strategy, budget is very important. Every brand wants to reach maximum target audience using all possible channels but

18 of 60

that is not possible as everything costs money and we need to optimize costs and hence the budget impacts the media strategy.

□ 3. Timing: Timing is an important aspect of media strategy. When to show the messaging to the customers can make all the difference.

The timing of advertisement is very critical especially with respect to the seasonal products.

□ 4. Channel: Channels and locations are quite similar in current context where online media is very relevant but for conventional advertising and messaging, a lot of channels like 1. TV; 2. Print; 3. Radio are still very relevant and used extensively in the media strategy.

4.10 Qualitative Research:

Qualitative market research is an open ended questions((conversational) based research method that heavily relies on the following market research methods:: focus groups, in-depth interviews, and other innovative research methods. It is based on a small but highly validated sample size, usually consisting of 6 to 10 respondents.

4.10.1 Qualitative research approaches

Approach

What does it involve?

Grounded theory

Researchers collect rich data on a topic of interest and develop theories inductively.

Researchers immerse themselves in groups or organizations to understand their cultures.

Action research

Researchers and participants collaboratively link theory to practice to drive social change.

Phenomenological

research

Researchers investigate a phenomenon or event by describing and interpreting participants' lived experiences.

Narrative research

Researchers examine how stories are told to understand how participants perceive and make sense of their experiences.

19 of 60

4.10.2 Qualitative research methods

Each of the research approaches involve using one or more data collection methods.

These are some of the most common qualitative methods:

• Observations: recording what you have seen, heard, or encountered in detailed field notes.

• Interviews: personally asking people questions in one-on-one conversations.

• Focus groups: asking questions and generating discussion among a group of people.

Surveys: distributing questionnaires with open-ended questions.

• Secondary research: collecting existing data in the form of texts, images, audio or video recordings, etc.

4.10.3 Qualitative data analysis

Qualitative data can take the form of texts, photos, videos and audio. For example,

you might be working with interview transcripts, survey responses, fieldnotes, or

recordings from natural settings.

Most types of qualitative data analysis share the same five steps:

1. Prepare and organize your data. This may mean transcribing interviews or typing up fieldnotes.

2. Review and explore your data. Examine the data for patterns or repeated ideas that emerge.

3. Develop a data coding system. Based on your initial ideas, establish a set of codes that you can apply to categorize your data.

4. Assign codes to the data. For example, in qualitative survey analysis, this may mean going through each participant's responses and tagging them with codes in a spreadsheet. As you go through your data, you can create new codes to add to your system if necessary.

5. Identify recurring themes. Link codes together into cohesive, overarching themes.

20 of 60

Advantages of qualitative research

Qualitative research often tries to preserve the voice and perspective of participants and can be adjusted as new research questions arise. Qualitative research is good for:

• Flexibility

The data collection and analysis process can be adapted as new ideas or patterns emerge. They are not rigidly decided beforehand.

• Natural settings

Data collection occurs in real-world contexts or in naturalistic ways.

• Meaningful insights

Detailed descriptions of people's experiences, feelings and perceptions can be used in designing, testing or improving systems or products.

• Generation of new ideas

Open-ended responses mean that researchers can uncover novel problems or opportunities that they wouldn't have thought of otherwise.

Disadvantages of qualitative research

Researchers must consider practical and theoretical limitations in analyzing and interpreting their data. Qualitative research suffers from:

• Unreliability

The real-world setting often makes qualitative research unreliable because of uncontrolled factors that affect the data.

• Subjectivity

Due to the researcher's primary role in analyzing and interpreting data, qualitative research cannot be replicated. The researcher decides what is important and what is irrelevant in data analysis, so interpretations of the same data can vary greatly.

• Limited generalizability

Small samples are often used to gather detailed data about specific contexts. Despite rigorous analysis procedures, it is difficult to draw generalizable conclusions because the data may be biased and unrepresentative of the wider population.

• Labor-intensive

Although software can be used to manage and record large amounts of text, data analysis often has to be checked or performed manually.

21 of 60

Unit IV – Multivariate analysis

  • Use of various statistical tools – Descriptive & Inference Statistics, Statistical Hypothesis Testing, Multivariate Analysis - Discriminant Analysis, Cluster Analysis, Segmenting and Positioning, Factor Analysis

22 of 60

Use of various statistical tools

  • Introduction: Market research relies heavily on statistical techniques in order to bring more insights to the usual deliverables and outputs. Analysing the collected data with basics tools is a fundamental aspect but sometimes a statistical methodology can answer the client’s question in a better way.
  • In the context of market research the researcher samples customers from populations to establish their perception towards particular products and services, or to identify purchasing behaviour so as to predict future preferences or buying habits.
  • The information gathered in these surveys can then be used to draw inference about the wider population with a certain level of statistical confidence that the results are accurate.
  • A necessary prerequisite to conducting a survey, and subsequently to drawing inference about a population, is to decide upon the best method of data collection.
  • Data collection encompasses the fundamental areas of survey design and sampling.
  • Analysing the collected data is another fundamental aspect and can include any number of statistical techniques.
  • A broad understanding of numerical data and an ability to interpret graphical and numerical descriptive measures is an important starting point for becoming proficient at data collection, analysis and interpretation of results.

In this chapter, statistical techniques commonly used in a market research environment to draw inference from survey data is discussed.

23 of 60

Descriptive statistics

Population Vs Sample

The image illustrates the concept of population and sample. Using random sample measurements from a representative group, we can estimate, predict, or infer characteristics about the larger population. While there are many technical variations on this technique, they all follow the same underlying principles.

Descriptive Statistics:

  • Descriptive statistics are brief descriptive coefficients that summarize a given data set, which can be either a representation of the entire population or a sample of a population.
  • The term ‘descriptive statistics’ can be used to describe both individual quantitative observations (also known as ‘summary statistics’) as well as the overall process of obtaining insights from these data.
  • Descriptive statistics may be used to describe both an entire population or an individual sample.
  • Because they are merely explanatory, descriptive statistics are not heavily concerned with the differences between the two types of data.

24 of 60

  • Descriptive statistics are broken down into measures of central tendency and measures of variability (spread).
  • Measures of central tendency include the mean, median, and mode, while measures of variability include standard deviation, variance, minimum and maximum variables, kurtosis, and skewness.

  • Descriptive statistics, describe and understand the features of a specific data set by giving short summaries about the sample and measures of the data.
  • The most recognized types of descriptive statistics are measures of center: the meanmedian, and mode, which are used at almost all levels of math and statistics. The mean, or the average, is calculated by adding all the figures within the data set and then dividing by the number of figures within the set.

I Measures of Central Tendency :

It is the middle point of a distribution. Tabulated data provides the data in a systematic order and enhances their understanding. Generally, in any distribution values of the variables tend to cluster around a central value of the distribution. This tendency of the distribution is known as central tendency and measures devised to consider this tendency is know as measures of central tendency. A measure of central tendency is useful if it represents accurately the distribution of scores on which it is based.

25 of 60

Characteristics of a good measure of central tendency :

  • It should be clearly defined- The definition of a measure of central tendency should be clear and unambiguous so that it leads to one and only one information.
  • It should be readily comprehensible and easy to compute.
  • It should be based on all observations- A good measure of central tendency should be based on all the values of the distribution of scores.
  • It should be amenable for further mathematical treatment.
  • It should be least affected by the fluctuation of sampling.

In statistics there are three most commonly used measures of central tendency., viz. Arithmetic Mean , Median, and Mode.

  1. Arithmetic Mean: The arithmetic mean is most popular and widely used measure of central tendency. This is obtained by dividing the sum of the values of the variable by the number of values. It is also a useful measure for further statistics and comparisons among different data sets.

One of the major limitations of arithmetic mean is that it cannot be computed for open-ended class-intervals.

2) Median: Median is the middle most value in a data distribution. It divides the distribution into two equal parts so that exactly one half of the observations is below and one half is above that point. Since median clearly denotes the position of an observation in an array, it is also called a position average. Thus more technically, median of an array of numbers arranged in order of their magnitude is either the middle value or the arithmetic mean of the two middle values. It is not affected by extreme values in the distribution.

3) Mode: Mode is the value in a distribution that corresponds to the maximum concentration of frequencies. It may be regarded as the most typical of a series value. In more simple words, mode is the point in the distribution comprising maximum frequencies therein

26 of 60

INFERENTIAL STATISTICS

  • Inferential statistics is a branch of statistics that makes the use of various analytical tools to draw inferences about the population data from sample data.  Inferential statistics help to draw conclusions about the population while descriptive statistics summarizes the features of the data set.
  • There are two main types of inferential statistics - hypothesis testing and regression analysis.
  • The samples chosen in inferential statistics need to be representative of the entire population.

Types of Inferential Statistics

  • Inferential statistics can be classified into hypothesis testing and regression analysis. Hypothesis testing also includes the use of confidence intervals to test the parameters of a population. Given below are the different types of inferential statistics.

Hypothesis Testing

  • Hypothesis testing is a type of inferential statistics that is used to test assumptions and draw conclusions about the population from the available sample data. It involves setting up a null hypothesis and an alternative hypothesis followed by conducting a statistical test of significance. A conclusion is drawn based on the value of the test statistic, the critical value, and the confidence intervals. A hypothesis test can be left-tailed, right-tailed, and two-tailed. 

27 of 60

Regression Analysis

  • Regression analysis is used to quantify how one variable will change with respect to another variable. There are many types of regressions available such as simple linear, multiple linear, nominal, logistic, and ordinal regression. The most commonly used regression in inferential statistics is linear regression. Linear regression checks the effect of a unit change of the independent variable in the dependent variable. 

Inferential Statistics vs Descriptive Statistics

  • Descriptive and inferential statistics are used to describe data and make generalizations about the population from samples.

Inferential Statistics

Descriptive Statistics

Inferential statistics are used to make conclusions about the population by using analytical tools on the sample data.

Descriptive statistics are used to quantify the characteristics of the data.

Hypothesis testing and regression analysis are the analytical tools used.

Measures of central tendency and measures of dispersion are the important tools used.

It is used to make inferences about an unknown population

It is used to describe the characteristics of a known sample or population.

Measures of inferential statistics are t-test, z test, linear regression, etc.

Measures of descriptive statistics are variance, range, mean, median, etc.

28 of 60

Hypothesis testing

  • When interpreting research findings, researchers need to assess whether these findings may have occurred by chance. Hypothesis testing is a systematic procedure for deciding whether the results of a research study support a particular theory which applies to a population.
  • Hypothesis testing uses sample data to evaluate a hypothesis about a population. A hypothesis test assesses how unusual the result is, whether it is reasonable chance variation or whether the result is too extreme to be considered chance variation.

Hypothesis

  • A hypothesis is a calculated prediction or assumption about a population parameter based on limited evidence. The whole idea behind hypothesis formulation is testing—this means the researcher subjects his or her calculated assumption to a series of evaluations to know whether they are true or false. 
  • A null hypothesis is a statement, in which there is no relationship between two variables. An alternative hypothesis is a statement; that is simply the inverse of the null hypothesis, i.e. there is some statistical significance between two measured phenomenon.

Stages of Hypothesis Testing

The five (5) stages of hypothesis testing are:

    • Determine the null hypothesis
    • Specify the alternative hypothesis
    • Set the significance level
    • Calculate the test statistics and corresponding P-value
    • Draw your conclusion

Determine the Null Hypothesis

  • Hypothesis testing starts with creating a null hypothesis which stands as an assumption that a certain statement is false or implausible. For example, the null hypothesis (H0) could suggest that different subgroups in the research population react to a variable in the same way. 

29 of 60

Specify the Alternative Hypothesis

  • The next step is to determine the alternative hypothesis. The alternative hypothesis counters the null assumption by suggesting the statement or assertion is true. Depending on the purpose of the research, the alternative hypothesis can be one-sided or two-sided. 

Set the Significance Level

  • Many researchers create a 5% allowance for accepting the value of an alternative hypothesis, even if the value is untrue. This means that there is a 0.05 chance that one would go with the value of the alternative hypothesis, despite the truth of the null hypothesis. 
  • Smaller the significance level, the greater the burden of proof needed to reject the null hypothesis and support the alternative hypothesis.

Calculate the Test Statistics and Corresponding P-Value 

  • Test statistics in hypothesis testing allows to compare different groups between variables while the p-value accounts for the probability of obtaining sample statistics if your null hypothesis is true. In this case, the test statistics can be the mean, median and similar parameters. 
  • If the p-value is 0.65, for example, then it means that the variable in the hypothesis will happen 65 in100 times by pure chance.

Draw Your Conclusions

  • After conducting a series of tests, the researcher will agree or refute the hypothesis based on feedback and insights from the sample data.  

30 of 60

Applications of Hypothesis Testing in Research

Hypothesis testing isn't only confined to numbers and calculations; it also has several real-life applications in business, manufacturing, advertising, and medicine. 

  • In a factory or other manufacturing plants, hypothesis testing is an important part of quality and production control before the final products are approved and sent out to the consumer. 
  • During ideation and strategy development, C-level executives use hypothesis testing to evaluate their theories and assumptions before any form of implementation. For example, they could leverage hypothesis testing to determine whether or not some new advertising campaign, marketing technique, etc. causes increased sales. 
  • In addition, hypothesis testing is used during clinical trials to prove the efficacy of a drug or new medical method before its approval for widespread human usage. 

Importance/Benefits of Hypothesis Testing 

Other benefits include: 

  • Hypothesis testing provides a reliable framework for making any data decisions for your population of interest. 
  • It helps the researcher to successfully extrapolate data from the sample to the larger population. 
  • Hypothesis testing allows the researcher to determine whether the data from the sample is statistically significant. 
  • Hypothesis testing is one of the most important processes for measuring the validity and reliability of outcomes in any systematic investigation. 
  • It helps to provide links to the underlying theory and specific research questions.

31 of 60

MULTI VARIATE ANALYSIS

Introduction

Multivariate means involving multiple dependent variables resulting in one outcome. This explains that the majority of the problems in the real world are Multivariate.  For example, we cannot predict the weather of any year based on the season. There are multiple factors like pollution, humidity, precipitation, etc. Here, we will introduce you to multivariate analysis, its history, and its application in different fields. 

Multivariate analysis (MVA) is a Statistical procedure for analysis of data involving more than one type of measurement or observation. It may also mean solving problems where more than one dependent variable is analyzed simultaneously with other variables.

Advantages of Multivariate Analysis

  • The main advantage of multivariate analysis is that since it considers more than one factor of independent variables that influence the variability of dependent variables, the conclusion drawn is more accurate.
  • The conclusions are more realistic and nearer to the real-life situation.

Disadvantages of Multivariate Analysis

  • The main disadvantage of MVA includes that it requires rather complex computations to arrive at a satisfactory conclusion.
  • Many observations for a large number of variables need to be collected and tabulated; it is a rather time-consuming process.

32 of 60

Classification of Multivariate Techniques

  • Multivariate analysis technique can be classified into two broad categories viz., This classification depends upon the question: are the involved variables dependent on each other or not? 
  • If the answer is yes:  Dependence methods.�If the answer is no:  Interdependence methods. 

Dependence technique:  Dependence Techniques are types of multivariate analysis techniques that are used when one or more of the variables can be identified as dependent variables and the remaining variables can be identified as independent.

Interdependence Technique

  • Interdependence techniques are a type of relationship that variables cannot be classified as either dependent or independent. 
  • It aims to unravel relationships between variables and/or subjects without explicitly assuming specific distributions for the variables. The idea is to describe the patterns in the data without making (very) strong assumptions about the variables. 

33 of 60

Factor Analysis 

  • Factor analysis is a way to condense the data in many variables into just a few variables.
  • It is also sometimes called “dimension reduction”.
  • It makes the grouping of variables with high correlation.
  • Factor analysis includes techniques such as principal component analysis and common factor analysis.
  • This type of technique is used as a pre-processing step to transform the data before using other models.
  • When the data has too many variables, the performance of multivariate techniques is not at the optimum level, as patterns are more difficult to find.
  • By using factor analysis, the patterns become less diluted and easier to analyze.

Objectives of factor analysis

  • To definitively understand how many factors are needed to explain common themes amongst a given set of variables.
  • To determine the extent to which each variable in the dataset is associated with a common theme or factor.
  • To provide an interpretation of the common factors in the dataset.
  • To determine the degree to which each observed data point represents each theme or factor.

Forms of Factor Analysis

  • Exploratory Factor Analysis should be used when you need to develop a hypothesis about a relationship between variables. 
  • Confirmatory Factor Analysis should be used to test a hypothesis about the relationship between variables.
  • Construct Validity should be used to test the degree to which your survey actually measures what it is intended to measure.

34 of 60

Assumptions:

  • No outlier: Assume that there are no outliers in data.
  • Adequate sample size: The case must be greater than the factor.
  • No perfect multicollinearity: Factor analysis is an interdependency technique.  There should not be perfect multicollinearity between the variables.
  • Homoscedasticity: Since factor analysis is a linear function of measured variables, it does not require homoscedasticity between the variables.
  • Linearity: Factor analysis is also based on linearity assumption.  Non-linear variables can also be used.  After transfer, however, it changes into linear variable.
  • Interval Data: Interval data are assumed.

Types of factoring:

There are different types of methods used to extract the factor from the data set:

1. Principal component analysis: This is the most common method used by researchers.  PCA starts extracting the maximum variance and puts them into the first factor.  After that, it removes that variance explained by the first factors and then starts extracting maximum variance for the second factor.  This process goes to the last factor.

2. Common factor analysis: The second most preferred method by researchers, it extracts the common variance and puts them into factors.  This method does not include the unique variance of all variables.  This method is used in SEM.

3. Image factoring: This method is based on correlation matrix.  OLS Regression method is used to predict the factor in image factoring.

4. Maximum likelihood method: This method also works on correlation metric but it uses maximum likelihood method to factor.

5. Other methods of factor analysis: Alfa factoring outweighs least squares.  Weight square is another regression based method which is used for factoring.

Factor loading:

Factor loading is basically the correlation coefficient for the variable and factor.  Factor loading shows the variance explained by the variable on that particular factor. 

35 of 60

  • Eigenvalues: Eigenvalues is also called characteristic roots.  Eigenvalues shows variance explained by that particular factor out of the total variance. 

  For example, if our first factor explains 68% variance out of the total, this means that 32% variance will be explained by the other factor.

  • Factor score: The factor score is also called the component score.  This score is of all row and columns, which can be used as an index of all variables and can be used for further analysis. 

Criteria for determining the number of factors: 

  • Eigenvalues is a good criteria for determining a factor. 
  • If Eigenvalues is greater than one, we should consider that a factor and if Eigenvalues is less than one, then we should not consider that a factor. 
  • According to the variance extraction rule, it should be more than 0.7.  If variance is less than 0.7, then we should not consider that a factor.
  • Rotation method: Rotation method makes it more reliable to understand the output.  Eigenvalues do not affect the rotation method, but the rotation method affects the Eigenvalues or percentage of variance extracted.  There are a number of rotation methods available: (1) No rotation method, (2) Varimax rotation method, (3) Quartimax rotation method, (4) Direct oblimin rotation method, and (5) Promax rotation method. 
  • Each of these can be easily selected in SPSS, and we can compare our variance explained by those particular methods.

STEP by STEP procedure

36 of 60

Cluster analysis

  • Cluster analysis is a statistical method used to group similar objects into respective categories. It can also be referred to as segmentation analysis, taxonomy analysis, or clustering.
  • The goal of performing a cluster analysis is to sort different objects or data points into groups in a manner that the degree of association between two objects is high if they belong to the same group, and low if they belong to different groups.
  • Cluster analysis differs from many other statistical methods due to the fact that it’s mostly used when researchers do not have an assumed principle or fact that they are using as the foundation of their research.
  • It doesn’t make any distinction between dependent and independent variables. Instead, cluster analysis is leveraged mostly to discover structures in data without providing an explanation or interpretation. 
  • Put simply, cluster analysis discovers structures in data without explaining why those structures exist. 
  • For example, when cluster analysis is performed as part of market research, specific groups can be identified within a population. The analysis of these groups can then determine how likely a population cluster is to purchase products or services. If these groups are defined clearly, a marketing team can then target varying cluster with tailored, targeted communication. 
  • Common Applications of Cluster Analysis 
  • Marketing
  • Marketers commonly use cluster analysis to develop market segments, which allow for better positioning of products and messaging.  company to better position itself, explore new markets, and development products that specific clusters find relevant and valuable.  

37 of 60

  • Insurance  
  • Insurance companies often leverage cluster analysis if there are a high number of claims in a given region. This enables them to learn exactly what is driving this increase in claims.  
  • Geology  
  • For cities on fault lines, geologists use cluster analysis to evaluate seismic risk and the potential weaknesses of earthquake-prone regions. By considering the results of this research, residents can do their best to prepare mitigate potential damage. 

The Benefits of Cluster Analysis

  • Clustering allows researchers to identify and define patterns between data elements. 
  • Revealing these patterns between data points helps to distinguish and outline structures which might not have been apparent before, but which give significant meaning to the data once they are discovered.
  • Once a clearly defined structure emerges from the dataset at hand, informed decision-making becomes much easier.

The Different Types of Cluster Analysis

There are three primary methods used to perform cluster analysis:  

  • Hierarchical Cluster : This is the most common method of clustering. It creates a series of models with cluster solutions from 1 (all cases in one cluster) to n (each case is an individual cluster). This approach also works with variables instead of cases.  
  • Finally, hierarchical cluster analysis can handle nominal, ordinal, and scale data. But, remember not to mix different levels of measurement into your study.

K-Means Cluster

  • This method is used to quickly cluster large datasets. Here, researchers define the number of clusters prior to performing the actual study. This approach is useful when testing different models with a different assumed number of clusters.

38 of 60

Two-Step Cluster

  • This method uses a cluster algorithm to identify groupings by performing pre-clustering first, and then performing hierarchical methods. Two-step clustering is best for handling larger datasets that would otherwise take too long a time to calculate with strictly hierarchical methods. 
  • Essentially, two-step cluster analysis is a combination of hierarchical and k-means cluster analysis. It can handle both scale and ordinal data, and it automatically selects the number of clusters.

39 of 60

Discriminant Analysis

  • Discriminant analysis is a technique that is used by the researcher to analyze the research data when the criterion or the dependent variable is categorical and the predictor or the independent variable is interval in nature. The term categorical variable means that the dependent variable is divided into a number of categories. 

Discriminant analysis is statistical technique used to classify observations into non-overlapping groups, based on scores on one or more quantitative predictor variables.

  • Objectives
    • Development of discriminant functions
    • Examination of whether significant differences exist among the groups, in terms of the predictor variables.
    • Determination of which predictor variables contribute to most of the intergroup differences
    • Evaluation of the accuracy of classification

40 of 60

DA involves the determination of a linear equation like regression that will predict which group the case belongs to.

The form of the equation or function is:

D= v1 X1+ v2 X2+ v3X3+ ... + viXi + a

  • Where D = discriminate function
  • v = the discriminant coefficient or weight for that variable
  • X = respondent’s score for that variable 2
  • a = a constant
  • i = the number of predictor variables

Applications (Examples)

  1. An educational researcher may want to investigate which variables discriminate between high school graduates who decide (1) to go to college, (2) to attend a trade or professional school, or (3) to seek no further training or education. For that purpose the researcher could collect data on numerous variables prior to students' graduation. After graduation, most students will naturally fall into one of the three categories.

Discriminant Analysis could then be used to determine which variable(s) are the best predictors of students' subsequent educational choice.

2. A medical researcher may record different variables relating to patients' backgrounds in order to learn which variables best predict whether a patient is likely to recover completely (group 1), partially (group 2), or not at all (group 3). A biologist could record different characteristics of similar types (groups) of flowers, and then perform a discriminant function analysis to determine the set of characteristics that allows for the best discrimination between the types.

41 of 60

Unit IV – Multivariate analysis

  • Use of various statistical tools – Descriptive & Inference Statistics, Statistical Hypothesis Testing, Multivariate Analysis - Discriminant Analysis, Cluster Analysis, Segmenting and Positioning, Factor Analysis

42 of 60

Use of various statistical tools

  • Introduction: Market research relies heavily on statistical techniques in order to bring more insights to the usual deliverables and outputs. Analysing the collected data with basics tools is a fundamental aspect but sometimes a statistical methodology can answer the client’s question in a better way.
  • In the context of market research the researcher samples customers from populations to establish their perception towards particular products and services, or to identify purchasing behaviour so as to predict future preferences or buying habits.
  • The information gathered in these surveys can then be used to draw inference about the wider population with a certain level of statistical confidence that the results are accurate.
  • A necessary prerequisite to conducting a survey, and subsequently to drawing inference about a population, is to decide upon the best method of data collection.
  • Data collection encompasses the fundamental areas of survey design and sampling.
  • Analysing the collected data is another fundamental aspect and can include any number of statistical techniques.
  • A broad understanding of numerical data and an ability to interpret graphical and numerical descriptive measures is an important starting point for becoming proficient at data collection, analysis and interpretation of results.

In this chapter, statistical techniques commonly used in a market research environment to draw inference from survey data is discussed.

43 of 60

Descriptive statistics

Population Vs Sample

The image illustrates the concept of population and sample. Using random sample measurements from a representative group, we can estimate, predict, or infer characteristics about the larger population. While there are many technical variations on this technique, they all follow the same underlying principles.

Descriptive Statistics:

  • Descriptive statistics are brief descriptive coefficients that summarize a given data set, which can be either a representation of the entire population or a sample of a population.
  • The term ‘descriptive statistics’ can be used to describe both individual quantitative observations (also known as ‘summary statistics’) as well as the overall process of obtaining insights from these data.
  • Descriptive statistics may be used to describe both an entire population or an individual sample.
  • Because they are merely explanatory, descriptive statistics are not heavily concerned with the differences between the two types of data.

44 of 60

  • Descriptive statistics are broken down into measures of central tendency and measures of variability (spread).
  • Measures of central tendency include the mean, median, and mode, while measures of variability include standard deviation, variance, minimum and maximum variables, kurtosis, and skewness.

  • Descriptive statistics, describe and understand the features of a specific data set by giving short summaries about the sample and measures of the data.
  • The most recognized types of descriptive statistics are measures of center: the meanmedian, and mode, which are used at almost all levels of math and statistics. The mean, or the average, is calculated by adding all the figures within the data set and then dividing by the number of figures within the set.

I Measures of Central Tendency :

It is the middle point of a distribution. Tabulated data provides the data in a systematic order and enhances their understanding. Generally, in any distribution values of the variables tend to cluster around a central value of the distribution. This tendency of the distribution is known as central tendency and measures devised to consider this tendency is know as measures of central tendency. A measure of central tendency is useful if it represents accurately the distribution of scores on which it is based.

45 of 60

Characteristics of a good measure of central tendency :

  • It should be clearly defined- The definition of a measure of central tendency should be clear and unambiguous so that it leads to one and only one information.
  • It should be readily comprehensible and easy to compute.
  • It should be based on all observations- A good measure of central tendency should be based on all the values of the distribution of scores.
  • It should be amenable for further mathematical treatment.
  • It should be least affected by the fluctuation of sampling.

In statistics there are three most commonly used measures of central tendency., viz. Arithmetic Mean , Median, and Mode.

  • Arithmetic Mean: The arithmetic mean is most popular and widely used measure of central tendency. This is obtained by dividing the sum of the values of the variable by the number of values. It is also a useful measure for further statistics and comparisons among different data sets.

One of the major limitations of arithmetic mean is that it cannot be computed for open-ended class-intervals.

2) Median: Median is the middle most value in a data distribution. It divides the distribution into two equal parts so that exactly one half of the observations is below and one half is above that point. Since median clearly denotes the position of an observation in an array, it is also called a position average. Thus more technically, median of an array of numbers arranged in order of their magnitude is either the middle value or the arithmetic mean of the two middle values. It is not affected by extreme values in the distribution.

3) Mode: Mode is the value in a distribution that corresponds to the maximum concentration of frequencies. It may be regarded as the most typical of a series value. In more simple words, mode is the point in the distribution comprising maximum frequencies therein

46 of 60

INFERENTIAL STATISTICS

  • Inferential statistics is a branch of statistics that makes the use of various analytical tools to draw inferences about the population data from sample data.  Inferential statistics help to draw conclusions about the population while descriptive statistics summarizes the features of the data set.
  • There are two main types of inferential statistics - hypothesis testing and regression analysis.
  • The samples chosen in inferential statistics need to be representative of the entire population.

Types of Inferential Statistics

  • Inferential statistics can be classified into hypothesis testing and regression analysis. Hypothesis testing also includes the use of confidence intervals to test the parameters of a population. Given below are the different types of inferential statistics.

Hypothesis Testing

  • Hypothesis testing is a type of inferential statistics that is used to test assumptions and draw conclusions about the population from the available sample data. It involves setting up a null hypothesis and an alternative hypothesis followed by conducting a statistical test of significance. A conclusion is drawn based on the value of the test statistic, the critical value, and the confidence intervals. A hypothesis test can be left-tailed, right-tailed, and two-tailed. 

47 of 60

Regression Analysis

  • Regression analysis is used to quantify how one variable will change with respect to another variable. There are many types of regressions available such as simple linear, multiple linear, nominal, logistic, and ordinal regression. The most commonly used regression in inferential statistics is linear regression. Linear regression checks the effect of a unit change of the independent variable in the dependent variable. 

Inferential Statistics vs Descriptive Statistics

  • Descriptive and inferential statistics are used to describe data and make generalizations about the population from samples.

Inferential Statistics

Descriptive Statistics

Inferential statistics are used to make conclusions about the population by using analytical tools on the sample data.

Descriptive statistics are used to quantify the characteristics of the data.

Hypothesis testing and regression analysis are the analytical tools used.

Measures of central tendency and measures of dispersion are the important tools used.

It is used to make inferences about an unknown population

It is used to describe the characteristics of a known sample or population.

Measures of inferential statistics are t-test, z test, linear regression, etc.

Measures of descriptive statistics are variance, range, mean, median, etc.

48 of 60

Hypothesis testing

  • When interpreting research findings, researchers need to assess whether these findings may have occurred by chance. Hypothesis testing is a systematic procedure for deciding whether the results of a research study support a particular theory which applies to a population.
  • Hypothesis testing uses sample data to evaluate a hypothesis about a population. A hypothesis test assesses how unusual the result is, whether it is reasonable chance variation or whether the result is too extreme to be considered chance variation.

Hypothesis

  • A hypothesis is a calculated prediction or assumption about a population parameter based on limited evidence. The whole idea behind hypothesis formulation is testing—this means the researcher subjects his or her calculated assumption to a series of evaluations to know whether they are true or false. 
  • A null hypothesis is a statement, in which there is no relationship between two variables. An alternative hypothesis is a statement; that is simply the inverse of the null hypothesis, i.e. there is some statistical significance between two measured phenomenon.

Stages of Hypothesis Testing

The five (5) stages of hypothesis testing are:

    • Determine the null hypothesis
    • Specify the alternative hypothesis
    • Set the significance level
    • Calculate the test statistics and corresponding P-value
    • Draw your conclusion

Determine the Null Hypothesis

  • Hypothesis testing starts with creating a null hypothesis which stands as an assumption that a certain statement is false or implausible. For example, the null hypothesis (H0) could suggest that different subgroups in the research population react to a variable in the same way. 

49 of 60

Specify the Alternative Hypothesis

  • The next step is to determine the alternative hypothesis. The alternative hypothesis counters the null assumption by suggesting the statement or assertion is true. Depending on the purpose of the research, the alternative hypothesis can be one-sided or two-sided. 

Set the Significance Level

  • Many researchers create a 5% allowance for accepting the value of an alternative hypothesis, even if the value is untrue. This means that there is a 0.05 chance that one would go with the value of the alternative hypothesis, despite the truth of the null hypothesis. 
  • Smaller the significance level, the greater the burden of proof needed to reject the null hypothesis and support the alternative hypothesis.

Calculate the Test Statistics and Corresponding P-Value 

  • Test statistics in hypothesis testing allows to compare different groups between variables while the p-value accounts for the probability of obtaining sample statistics if your null hypothesis is true. In this case, the test statistics can be the mean, median and similar parameters. 
  • If the p-value is 0.65, for example, then it means that the variable in the hypothesis will happen 65 in100 times by pure chance.

Draw Your Conclusions

  • After conducting a series of tests, the researcher will agree or refute the hypothesis based on feedback and insights from the sample data.  

50 of 60

Applications of Hypothesis Testing in Research

Hypothesis testing isn't only confined to numbers and calculations; it also has several real-life applications in business, manufacturing, advertising, and medicine. 

  • In a factory or other manufacturing plants, hypothesis testing is an important part of quality and production control before the final products are approved and sent out to the consumer. 
  • During ideation and strategy development, C-level executives use hypothesis testing to evaluate their theories and assumptions before any form of implementation. For example, they could leverage hypothesis testing to determine whether or not some new advertising campaign, marketing technique, etc. causes increased sales. 
  • In addition, hypothesis testing is used during clinical trials to prove the efficacy of a drug or new medical method before its approval for widespread human usage. 

Importance/Benefits of Hypothesis Testing 

Other benefits include: 

  • Hypothesis testing provides a reliable framework for making any data decisions for your population of interest. 
  • It helps the researcher to successfully extrapolate data from the sample to the larger population. 
  • Hypothesis testing allows the researcher to determine whether the data from the sample is statistically significant. 
  • Hypothesis testing is one of the most important processes for measuring the validity and reliability of outcomes in any systematic investigation. 
  • It helps to provide links to the underlying theory and specific research questions.

51 of 60

MULTI VARIATE ANALYSIS

Introduction

Multivariate means involving multiple dependent variables resulting in one outcome. This explains that the majority of the problems in the real world are Multivariate.  For example, we cannot predict the weather of any year based on the season. There are multiple factors like pollution, humidity, precipitation, etc. Here, we will introduce you to multivariate analysis, its history, and its application in different fields. 

Multivariate analysis (MVA) is a Statistical procedure for analysis of data involving more than one type of measurement or observation. It may also mean solving problems where more than one dependent variable is analyzed simultaneously with other variables.

Advantages of Multivariate Analysis

  • The main advantage of multivariate analysis is that since it considers more than one factor of independent variables that influence the variability of dependent variables, the conclusion drawn is more accurate.
  • The conclusions are more realistic and nearer to the real-life situation.

Disadvantages of Multivariate Analysis

  • The main disadvantage of MVA includes that it requires rather complex computations to arrive at a satisfactory conclusion.
  • Many observations for a large number of variables need to be collected and tabulated; it is a rather time-consuming process.

52 of 60

Classification of Multivariate Techniques

  • Multivariate analysis technique can be classified into two broad categories viz., This classification depends upon the question: are the involved variables dependent on each other or not? 
  • If the answer is yes:  Dependence methods.�If the answer is no:  Interdependence methods. 

Dependence technique:  Dependence Techniques are types of multivariate analysis techniques that are used when one or more of the variables can be identified as dependent variables and the remaining variables can be identified as independent.

Interdependence Technique

  • Interdependence techniques are a type of relationship that variables cannot be classified as either dependent or independent. 
  • It aims to unravel relationships between variables and/or subjects without explicitly assuming specific distributions for the variables. The idea is to describe the patterns in the data without making (very) strong assumptions about the variables. 

53 of 60

Factor Analysis 

  • Factor analysis is a way to condense the data in many variables into just a few variables.
  • It is also sometimes called “dimension reduction”.
  • It makes the grouping of variables with high correlation.
  • Factor analysis includes techniques such as principal component analysis and common factor analysis.
  • This type of technique is used as a pre-processing step to transform the data before using other models.
  • When the data has too many variables, the performance of multivariate techniques is not at the optimum level, as patterns are more difficult to find.
  • By using factor analysis, the patterns become less diluted and easier to analyze.

Objectives of factor analysis

  • To definitively understand how many factors are needed to explain common themes amongst a given set of variables.
  • To determine the extent to which each variable in the dataset is associated with a common theme or factor.
  • To provide an interpretation of the common factors in the dataset.
  • To determine the degree to which each observed data point represents each theme or factor.

Forms of Factor Analysis

  • Exploratory Factor Analysis should be used when you need to develop a hypothesis about a relationship between variables. 
  • Confirmatory Factor Analysis should be used to test a hypothesis about the relationship between variables.
  • Construct Validity should be used to test the degree to which your survey actually measures what it is intended to measure.

54 of 60

Assumptions:

  • No outlier: Assume that there are no outliers in data.
  • Adequate sample size: The case must be greater than the factor.
  • No perfect multicollinearity: Factor analysis is an interdependency technique.  There should not be perfect multicollinearity between the variables.
  • Homoscedasticity: Since factor analysis is a linear function of measured variables, it does not require homoscedasticity between the variables.
  • Linearity: Factor analysis is also based on linearity assumption.  Non-linear variables can also be used.  After transfer, however, it changes into linear variable.
  • Interval Data: Interval data are assumed.

Types of factoring:

There are different types of methods used to extract the factor from the data set:

1. Principal component analysis: This is the most common method used by researchers.  PCA starts extracting the maximum variance and puts them into the first factor.  After that, it removes that variance explained by the first factors and then starts extracting maximum variance for the second factor.  This process goes to the last factor.

2. Common factor analysis: The second most preferred method by researchers, it extracts the common variance and puts them into factors.  This method does not include the unique variance of all variables.  This method is used in SEM.

3. Image factoring: This method is based on correlation matrix.  OLS Regression method is used to predict the factor in image factoring.

4. Maximum likelihood method: This method also works on correlation metric but it uses maximum likelihood method to factor.

5. Other methods of factor analysis: Alfa factoring outweighs least squares.  Weight square is another regression based method which is used for factoring.

Factor loading:

Factor loading is basically the correlation coefficient for the variable and factor.  Factor loading shows the variance explained by the variable on that particular factor. 

55 of 60

  • Eigenvalues: Eigenvalues is also called characteristic roots.  Eigenvalues shows variance explained by that particular factor out of the total variance. 

  For example, if our first factor explains 68% variance out of the total, this means that 32% variance will be explained by the other factor.

  • Factor score: The factor score is also called the component score.  This score is of all row and columns, which can be used as an index of all variables and can be used for further analysis. 

Criteria for determining the number of factors: 

  • Eigenvalues is a good criteria for determining a factor. 
  • If Eigenvalues is greater than one, we should consider that a factor and if Eigenvalues is less than one, then we should not consider that a factor. 
  • According to the variance extraction rule, it should be more than 0.7.  If variance is less than 0.7, then we should not consider that a factor.
  • Rotation method: Rotation method makes it more reliable to understand the output.  Eigenvalues do not affect the rotation method, but the rotation method affects the Eigenvalues or percentage of variance extracted.  There are a number of rotation methods available: (1) No rotation method, (2) Varimax rotation method, (3) Quartimax rotation method, (4) Direct oblimin rotation method, and (5) Promax rotation method. 
  • Each of these can be easily selected in SPSS, and we can compare our variance explained by those particular methods.

STEP by STEP procedure

56 of 60

Cluster analysis

  • Cluster analysis is a statistical method used to group similar objects into respective categories. It can also be referred to as segmentation analysis, taxonomy analysis, or clustering.
  • The goal of performing a cluster analysis is to sort different objects or data points into groups in a manner that the degree of association between two objects is high if they belong to the same group, and low if they belong to different groups.
  • Cluster analysis differs from many other statistical methods due to the fact that it’s mostly used when researchers do not have an assumed principle or fact that they are using as the foundation of their research.
  • It doesn’t make any distinction between dependent and independent variables. Instead, cluster analysis is leveraged mostly to discover structures in data without providing an explanation or interpretation. 
  • Put simply, cluster analysis discovers structures in data without explaining why those structures exist. 
  • For example, when cluster analysis is performed as part of market research, specific groups can be identified within a population. The analysis of these groups can then determine how likely a population cluster is to purchase products or services. If these groups are defined clearly, a marketing team can then target varying cluster with tailored, targeted communication. 
  • Common Applications of Cluster Analysis 
  • Marketing
  • Marketers commonly use cluster analysis to develop market segments, which allow for better positioning of products and messaging.  company to better position itself, explore new markets, and development products that specific clusters find relevant and valuable.  

57 of 60

  • Insurance  
  • Insurance companies often leverage cluster analysis if there are a high number of claims in a given region. This enables them to learn exactly what is driving this increase in claims.  
  • Geology  
  • For cities on fault lines, geologists use cluster analysis to evaluate seismic risk and the potential weaknesses of earthquake-prone regions. By considering the results of this research, residents can do their best to prepare mitigate potential damage. 

The Benefits of Cluster Analysis

  • Clustering allows researchers to identify and define patterns between data elements. 
  • Revealing these patterns between data points helps to distinguish and outline structures which might not have been apparent before, but which give significant meaning to the data once they are discovered.
  • Once a clearly defined structure emerges from the dataset at hand, informed decision-making becomes much easier.

The Different Types of Cluster Analysis

There are three primary methods used to perform cluster analysis:  

  • Hierarchical Cluster : This is the most common method of clustering. It creates a series of models with cluster solutions from 1 (all cases in one cluster) to n (each case is an individual cluster). This approach also works with variables instead of cases.  
  • Finally, hierarchical cluster analysis can handle nominal, ordinal, and scale data. But, remember not to mix different levels of measurement into your study.

K-Means Cluster

  • This method is used to quickly cluster large datasets. Here, researchers define the number of clusters prior to performing the actual study. This approach is useful when testing different models with a different assumed number of clusters.

58 of 60

Two-Step Cluster

  • This method uses a cluster algorithm to identify groupings by performing pre-clustering first, and then performing hierarchical methods. Two-step clustering is best for handling larger datasets that would otherwise take too long a time to calculate with strictly hierarchical methods. 
  • Essentially, two-step cluster analysis is a combination of hierarchical and k-means cluster analysis. It can handle both scale and ordinal data, and it automatically selects the number of clusters.

59 of 60

Discriminant Analysis

  • Discriminant analysis is a technique that is used by the researcher to analyze the research data when the criterion or the dependent variable is categorical and the predictor or the independent variable is interval in nature. The term categorical variable means that the dependent variable is divided into a number of categories. 

Discriminant analysis is statistical technique used to classify observations into non-overlapping groups, based on scores on one or more quantitative predictor variables.

  • Objectives
    • Development of discriminant functions
    • Examination of whether significant differences exist among the groups, in terms of the predictor variables.
    • Determination of which predictor variables contribute to most of the intergroup differences
    • Evaluation of the accuracy of classification

60 of 60

DA involves the determination of a linear equation like regression that will predict which group the case belongs to.

The form of the equation or function is:

D= v1 X1+ v2 X2+ v3X3+ ... + viXi + a

  • Where D = discriminate function
  • v = the discriminant coefficient or weight for that variable
  • X = respondent’s score for that variable 2
  • a = a constant
  • i = the number of predictor variables

Applications (Examples)

  • An educational researcher may want to investigate which variables discriminate between high school graduates who decide (1) to go to college, (2) to attend a trade or professional school, or (3) to seek no further training or education. For that purpose the researcher could collect data on numerous variables prior to students' graduation. After graduation, most students will naturally fall into one of the three categories.

Discriminant Analysis could then be used to determine which variable(s) are the best predictors of students' subsequent educational choice.

2. A medical researcher may record different variables relating to patients' backgrounds in order to learn which variables best predict whether a patient is likely to recover completely (group 1), partially (group 2), or not at all (group 3). A biologist could record different characteristics of similar types (groups) of flowers, and then perform a discriminant function analysis to determine the set of characteristics that allows for the best discrimination between the types.