Five steps to prepare quantitative data for analysis

Analysis is the process of translating your collected data into something recognisable and understandable. You start by getting the data ready for inspection. Students often ask us to show them the five steps to prepare quantitative data for analysis.

Take Lorcan Ball, who recently completed his MA in Tourism Studies. For his capstone project, Lorcan investigated the relationship between the size of a household and how much money they spent on an overseas vacation. His central hypothesis was that as household size rises, overall vacation expenditure increases, both absolutely and relatively.

Lorcan employed a survey strategy and a quantitative methodology to gather measurable data. He distributed hard-copy, researcher-completed questionnaires at airports and ferry terminals. He used simple stratified random sampling based on household size to select his sample. All participants were asked the same questions from lists of pre-set options. Then Lorcan rated and ranked their answers against pre-determined Likert scales.

We reminded Lorcan that data preparation for quantitative analysis involves converting the data into clean, numerical formats. One of our Directors, Dr. Sue Mulhall, helped him prepare his data in five systematic steps.

Step 1 – Cleansing

Lorcan cleansed the data by identifying any errors and then normalising, or excluding, them. Even though his questionnaire contained 10 questions, the number ‘11’ was mistakenly used for one of the question numbers. He asked Sue what to do with that question. On checking, it was evident that question 8 had been mislabelled as question 11. Lorcan changed the erroneous number (11) to the correct one (8). Sue used this experience as a learning occasion. She explained that sometimes a question may have to be omitted if the labelling error is not so obvious to pinpoint or easy to resolve.

Step 2 – Coding

Lorcan coded the data in two phases.

First, allocating a numerical code to each of the 10 questions. These covered areas such as household size, expenditure, destination, duration, activities and accommodation. The first question referred to household size. Thus, it was assigned the code ‘1’. The second question related to expenditure, so was allotted the code ‘2’, and so on.

Second, allocating a variable to each response for every question. For example, the household size details from question one were categorised as follows: one adult = ‘1’; two adults = ‘2’; one adult with children = ‘3’; and two adults with children = ‘4’. Variables ‘3’ and ‘4’ had additional sub-categories. This was because the questionnaires accumulated information about number of children, and their status and respective ages. To illustrate: one adult with one dependent child = 3.1; and one adult with two dependent children = 3.2, and so on.

Step 3 – Framing

Lorcan created a coding frame using a well-known, proprietary spreadsheet software package. He built a database with suitable headings for the various columns and rows. Each column heading designated the allocated numerical code for every one of the 10 questions. Each row heading assigned the allocated variable for every one of the set responses to all 10 questions.

Step 4 – Transferring

Lorcan transferred the clean, coded data from the questionnaires into the coding frame. He inputted the relevant data into their specific placeholders on the spreadsheet. The 10 question codes were entered into columns ‘A’ to ‘J’ (one question per column). The response variables were then entered into their requisite rows. For instance, question one about household size, which was assigned code ‘1’, consisted of four rows (A1, A2, A3, A4). One for each of the four household size alternatives available.

Step 5 – Reviewing

Lorcan reviewed the data in overview to acquire a ‘big picture’. He printed off the entire spreadsheet and pinned it on to a whiteboard. Then he identified, with Sue’s assistance, significant data sets and detected high-level associations between them.

Now that he had methodically prepared his data, Lorcan was ready for the next two phases of the analytical process. Using formulae to analyse the prepared data to generate his results and then applying statistical operations to interpret his results to produce his findings.

Leave a Reply

Your email address will not be published. Required fields are marked *