Introduction to Data and Data Collection
Definition of Data: Data consists of facts, numbers, letters, and symbols that describe an object, idea, condition, or situation.
Data Collection: This is the process of gathering and measuring information on variables of interest in a systematic fashion. It allows researchers to answer research questions, test hypotheses, and evaluate outcomes.
- Raw data refers to the initial information gathered (e.g., asking for a name) before it is processed, analyzed, or interpreted.
- It is a crucial step in both qualitative and quantitative research; incorrect data collection compromises the entire research study.
Sources of Data
Data sources are primarily divided into two categories: Primary Data and Secondary Data.
Primary Data (Primary Source)
- Definition: Data collected directly from the original source or first-hand information. The researcher collects this data themselves.
- Examples: Interviews, observations, questionnaires, and schedules.
Secondary Data (Secondary Source)
- Definition: Data that has already been collected by someone else for another purpose. It is indirect information.
- Characteristics: It is cost-effective and time-saving because the data already exists.
- Examples: Patient records (files), government reports, censuses, historical records, biographies, newspapers, and published/unpublished theses.
Types of Data
Data is broadly classified into Qualitative and Quantitative types.
Qualitative Data (Categorical Data)
- Definition: Non-numerical data regarding qualities, opinions, attitudes, values, and images.
- Sub-types of Qualitative Data:
- Nominal Data: Derived from the Latin word Nomen (meaning name). It involves naming or labelling variables without any quantitative value or order (e.g., Gender, Hair colour, Marital status, Introvert/Extrovert).
- Ordinal Data: Categorical data that follows a specific order, hierarchy, or rank, though the distance between them is not defined (e.g., First/Second/Third, Socio-economic status: Lower/Middle/Upper, Likert scales: Agree/Disagree).
- Methods of Collection: Interviews, focus groups, participant observation, and document analysis.
Quantitative Data (Numerical Data)
- Definition: Data dealing with numbers, quantities, amounts, and ranges. It answers questions like "How much?" or "How many?".
- Sub-types of Quantitative Data:
- Discrete Data: Finite, fixed numbers that cannot be subdivided into decimals. They are counted in whole numbers (e.g., Number of patients in a ward, Number of students, Number of questions answered correctly). Represented often by bar graphs.
- Continuous Data: Data that can change over time and includes infinite values, fractions, and decimals (e.g., Weight, Height, Temperature, Speed of a vehicle). Represented by histograms and line graphs.
Methods of Data Collection
Quantitative Methods
- Surveys and Questionnaires:
- Use close-ended questions (Yes/No, specific options).
- Can be self-administered, sent via email/phone, or conducted online to reach large numbers of respondents.
- Experiments:
- Used in controlled studies (scientific research) to determine cause-and-effect relationships where variables are manipulated.
Qualitative Methods
- Interviews:
- Involve in-depth, open-ended questions to explore complex behaviours and experiences.
- Structure of Interviews:
- Structured: Uses a pre-determined, standard set of questions with no variation.
- Semi-Structured: Uses a guide/key questions but allows the interviewer to diverge and explore ideas further.
- Unstructured: Totally unplanned, informal, and conversational; useful for exploring new topics without preconceived theories.
- Focus Groups:
- A group discussion (typically 6-12 people) guided by a facilitator (team leader) to gather collective views and insights on group dynamics.
- Observation:
- Systematic recording of behaviour or events in a natural setting.
- Types:
- Non-participant/Passive: The researcher observes without getting involved.
- Participant: The researcher enters the environment and interacts with participants to gain a deeper understanding.
Mixed Methods
- Triangulation: The process of combining multiple data collection methods (e.g., using surveys, interviews, and focus groups together) to validate results and provide a comprehensive understanding of the research problem.
Other Research Designs
- Longitudinal Study: Data is collected from the same subjects over a long period to study changes and development over time.
- Secondary Data Analysis: Analyzing existing data (like government reports) for a new purpose.
Data Analysis Software
- SPSS (Statistical Package for the Social Sciences): A software package used for quantitative research. It performs statistical analysis, data management, and data visualization (graphs).
- NVivo: A software tool used for qualitative research. It helps organize, manage, code, and analyze non-numerical data like text, audio, video, and images to identify themes and patterns.
Training Data Collectors
Training is vital to ensure data accuracy, consistency, and reliability.
Steps for Training:
- Orientation: Introduction to the study, objectives, and methodology.
- Detailed Protocol: Step-by-step instructions on data collection techniques.
- Role Play: Practicing interviews or data collection in simulated scenarios to refine skills.
- Supervision: Initial data collection should be supervised to ensure accuracy before independent collection begins.
Ethical Considerations in Data Collection
When collecting data, the following ethical principles must be followed:
- Informed Consent: Participants must understand the study and agree to participate voluntarily.
- Confidentiality: Protecting the personal information and identity of the participants.
- Minimise Harm: Ensuring the data collection process does not physically or emotionally harm the participants.