Introduction to Data and Data Collection

Definition of Data: Data consists of facts, numbers, letters, and symbols that describe an object, idea, condition, or situation.

Data Collection: This is the process of gathering and measuring information on variables of interest in a systematic fashion. It allows researchers to answer research questions, test hypotheses, and evaluate outcomes.

  • Raw data refers to the initial information gathered (e.g., asking for a name) before it is processed, analyzed, or interpreted.
  • It is a crucial step in both qualitative and quantitative research; incorrect data collection compromises the entire research study.

Sources of Data

Data sources are primarily divided into two categories: Primary Data and Secondary Data.

Primary Data (Primary Source)

  • Definition: Data collected directly from the original source or first-hand information. The researcher collects this data themselves.
  • Examples: Interviews, observations, questionnaires, and schedules.

Secondary Data (Secondary Source)

  • Definition: Data that has already been collected by someone else for another purpose. It is indirect information.
  • Characteristics: It is cost-effective and time-saving because the data already exists.
  • Examples: Patient records (files), government reports, censuses, historical records, biographies, newspapers, and published/unpublished theses.

Types of Data

Data is broadly classified into Qualitative and Quantitative types.

Qualitative Data (Categorical Data)

  • Definition: Non-numerical data regarding qualities, opinions, attitudes, values, and images.
  • Sub-types of Qualitative Data:
    1. Nominal Data: Derived from the Latin word Nomen (meaning name). It involves naming or labelling variables without any quantitative value or order (e.g., Gender, Hair colour, Marital status, Introvert/Extrovert).
    2. Ordinal Data: Categorical data that follows a specific order, hierarchy, or rank, though the distance between them is not defined (e.g., First/Second/Third, Socio-economic status: Lower/Middle/Upper, Likert scales: Agree/Disagree).
  • Methods of Collection: Interviews, focus groups, participant observation, and document analysis.

Quantitative Data (Numerical Data)

  • Definition: Data dealing with numbers, quantities, amounts, and ranges. It answers questions like "How much?" or "How many?".
  • Sub-types of Quantitative Data:
    1. Discrete Data: Finite, fixed numbers that cannot be subdivided into decimals. They are counted in whole numbers (e.g., Number of patients in a ward, Number of students, Number of questions answered correctly). Represented often by bar graphs.
    2. Continuous Data: Data that can change over time and includes infinite values, fractions, and decimals (e.g., Weight, Height, Temperature, Speed of a vehicle). Represented by histograms and line graphs.

Methods of Data Collection

Quantitative Methods

  • Surveys and Questionnaires:
    • Use close-ended questions (Yes/No, specific options).
    • Can be self-administered, sent via email/phone, or conducted online to reach large numbers of respondents.
  • Experiments:
    • Used in controlled studies (scientific research) to determine cause-and-effect relationships where variables are manipulated.

Qualitative Methods

  • Interviews:
    • Involve in-depth, open-ended questions to explore complex behaviours and experiences.
    • Structure of Interviews:
      • Structured: Uses a pre-determined, standard set of questions with no variation.
      • Semi-Structured: Uses a guide/key questions but allows the interviewer to diverge and explore ideas further.
      • Unstructured: Totally unplanned, informal, and conversational; useful for exploring new topics without preconceived theories.
  • Focus Groups:
    • A group discussion (typically 6-12 people) guided by a facilitator (team leader) to gather collective views and insights on group dynamics.
  • Observation:
    • Systematic recording of behaviour or events in a natural setting.
    • Types:
      • Non-participant/Passive: The researcher observes without getting involved.
      • Participant: The researcher enters the environment and interacts with participants to gain a deeper understanding.

Mixed Methods

  • Triangulation: The process of combining multiple data collection methods (e.g., using surveys, interviews, and focus groups together) to validate results and provide a comprehensive understanding of the research problem.

Other Research Designs

  • Longitudinal Study: Data is collected from the same subjects over a long period to study changes and development over time.
  • Secondary Data Analysis: Analyzing existing data (like government reports) for a new purpose.

Data Analysis Software

  • SPSS (Statistical Package for the Social Sciences): A software package used for quantitative research. It performs statistical analysis, data management, and data visualization (graphs).
  • NVivo: A software tool used for qualitative research. It helps organize, manage, code, and analyze non-numerical data like text, audio, video, and images to identify themes and patterns.

Training Data Collectors

Training is vital to ensure data accuracy, consistency, and reliability.

Steps for Training:

  1. Orientation: Introduction to the study, objectives, and methodology.
  2. Detailed Protocol: Step-by-step instructions on data collection techniques.
  3. Role Play: Practicing interviews or data collection in simulated scenarios to refine skills.
  4. Supervision: Initial data collection should be supervised to ensure accuracy before independent collection begins.

Ethical Considerations in Data Collection

When collecting data, the following ethical principles must be followed:

  1. Informed Consent: Participants must understand the study and agree to participate voluntarily.
  2. Confidentiality: Protecting the personal information and identity of the participants.
  3. Minimise Harm: Ensuring the data collection process does not physically or emotionally harm the participants.