2 Data and Its Types
Data is the raw material of every analysis you will ever run. Before you can compute a mean, draw a chart, or train a model, you need to know exactly what kind of data you are holding, because the type of data determines which summaries make sense, which charts are appropriate, and which statistical tests are even valid.
Most textbooks stop after two or three ways of classifying data. In practice a working analyst classifies every column along eight separate dimensions at once, and each one answers a different question about what that column will let you do. This chapter builds that full vocabulary, from where data comes from, through how it is measured, to how it is stored and how fast it arrives.
2.1 What Is Data?
Data refers to facts, figures, and observations collected about people, objects, or events, recorded in a form that can be communicated, stored, and analyzed. A single data point, such as one customer’s age or one product’s price, is called an observation. A collection of observations across one or more variables forms a dataset.
- Data: Raw, unprocessed facts, such as the number 42 or the word “Delhi” recorded against a customer record.
- Information: Data that has been organized and given context, such as “average customer age is 42 years.”
- Variable: A characteristic that can take different values across observations, such as age, income, or product category.
- Unit of analysis: The entity each row describes, such as a customer, a transaction, a store, or a month. Getting this wrong is the single most common structural error in applied analysis.
- Metadata: Data about the data, such as when a column was collected, in what unit it is measured, and what its permitted values are. Metadata is what makes a dataset reusable by someone who did not collect it.
Data on its own does not answer questions. A single number like 42 tells you nothing until it is organized, compared, and interpreted, which is precisely what the rest of this book teaches you to do, starting with the classification and tabulation methods in the next chapter.
2.2 Eight Ways to Classify the Same Column
A classification criterion is a question you ask about a variable. Different questions produce different answers, and the answers do not compete with one another, they stack. A single column such as monthly_income is simultaneously secondary in source, quantitative in nature, continuous in countability, ratio-scaled in measurement, cross-sectional in time reference, structured in form, ungrouped in organisation, and small-data in scale.
Learning to run all eight questions over every column, quickly and automatically, is what separates someone who has a dataset from someone who understands one.
2.3 1. Classification by Source
The first question to ask of any dataset is who collected it, and why. This determines how much you can trust it and how much cleaning it will need.
- Primary Data: Data collected firsthand by the analyst or organization for a specific purpose, through surveys, experiments, interviews, or direct observation. It is original and tailored to the research question, but collecting it costs time and money.
- Secondary Data: Data that already exists, having been collected by someone else for a different purpose, such as government statistics, company records, or published research. It is faster and cheaper to obtain, but may not perfectly fit the current question and its quality depends on the original collector.
| Aspect | Primary Data | Secondary Data |
|---|---|---|
| Collection purpose | The current research question | Some other, earlier purpose |
| Cost | High | Low to none |
| Time to obtain | Weeks to months | Hours to days |
| Control over quality | Full | None |
| Fit to the question | Exact | Approximate |
| Typical use | Testing a specific hypothesis | Framing a problem, benchmarking, context |
Secondary data is itself a large family, and marketing research texts classify it further by where it sits relative to the organisation. Internal secondary data already exists inside the firm; external secondary data comes from outside it.
-
Secondary Data
-
Internal
- Ready to useSales ledgers, CRM exports, billing records
- Requires further processingRaw server logs, scanned invoices, free-text tickets
-
External
-
Published materials
- General business sourcesGuides, directories, indexes, trade statistics
- Government sourcesCensus of India, NSSO rounds, RBI bulletins
-
Computerised databases
- Bibliographic, numeric, full-text, directoryOnline, internet, and offline databases
-
Syndicated services
- From householdsSurveys, purchase panels, media panels, scanner services
- From institutionsRetailer and wholesaler audits, industrial firm reports
-
Published materials
-
Internal
Example
A retail chain wants to understand why sales dropped in a region. Sending a fresh customer satisfaction survey to shoppers in that region produces primary data. Pulling last year’s point-of-sale transaction records from the company’s own database, originally logged for accounting purposes, produces internal secondary data. Buying a syndicated retail audit that reports category-level market share across competing chains produces external secondary data from institutions.
Secondary data is convenient, but never assume it is clean or complete simply because it comes from an official-looking source. Before relying on it, check six things: the specifications under which it was collected (sampling method, response rate, questionnaire wording), the error it may carry, its currency (how old is it?), the objective for which it was originally gathered, its nature (exactly how were the variables defined?), and its dependability (who published it, and what is their incentive?). A dataset that fails any one of these can quietly invalidate an entire analysis.
2.4 2. Classification by Nature
Every variable in a dataset falls into one of two broad families based on the nature of the values it can take.
- Qualitative (Categorical) Data: Describes a quality or category rather than a quantity. Values are labels, such as gender, city, or product type, and cannot be meaningfully averaged.
- Quantitative (Numerical) Data: Describes a measurable quantity expressed as a number, such as height, income, or number of orders, and supports arithmetic operations like addition and averaging.
Qualitative data is often subdivided by how many categories it admits:
- Dichotomous (binary): Exactly two categories, such as Yes/No, Pass/Fail, Churned/Retained. Binary variables are unusually convenient, because coding them 0 and 1 makes their mean a proportion.
- Polytomous (multichotomous): Three or more categories, such as state, payment method, or product line.
The 0/1 trick is worth internalising early. If churned is coded 1 for customers who left and 0 for those who stayed, then the mean of that column is exactly the churn rate. This is why binary variables appear everywhere in analytics, and why so many models are built to predict them.
Numbers stored in a column do not make it quantitative. A pin code of 600001, an employee ID of 4471, and a jersey number of 10 are all recorded as digits but are pure labels: averaging them is meaningless. The test is not “does it look like a number?” but “does arithmetic on it produce something that means anything?”
2.5 3. Classification by Countability
Quantitative data splits further according to whether its values can be counted or must be measured.
- Discrete Data: Takes only specific, countable values, usually whole numbers, with no meaningful values in between. The number of children in a household or the number of defective items in a batch are discrete, since you cannot have 2.5 children or 3.7 defects.
- Continuous Data: Can take any value within a range, including fractions and decimals, limited only by the precision of measurement. Height, weight, temperature, and time are continuous, since a person’s height could be 165.3 cm or 165.34 cm depending on how precisely it is measured.
Example
Consider a student records dataset with the columns: Name (Ravi, Priya, Aman), Gender (Male, Female), Marks Scored (78, 85, 91), and Number of Siblings (0, 1, 2). Name and Gender are qualitative, and Gender is specifically dichotomous. Marks Scored is quantitative and continuous, since a mark could in principle be 78.5. Number of Siblings is quantitative and discrete, since siblings are always counted in whole numbers.
Continuous data is almost always recorded discretely. Age is continuous in principle, but a survey that asks for age in completed years stores it as whole numbers. This does not make age a discrete variable, it makes it a continuous variable measured coarsely. The distinction matters because it decides whether a histogram (continuous) or a bar chart (discrete) is the honest picture, a choice taken up properly in Chapter 4.
2.6 4. Classification by Measurement Scale
Beyond qualitative and quantitative, every variable also has a level of measurement, sometimes called its measurement scale, which determines exactly what mathematical and statistical operations are valid on it. The four-level scheme below was set out by the psychologist S. S. Stevens in 1946 and remains the standard vocabulary across the sciences. Each level adds a capability the previous level lacked.
- Nominal: Categories with no inherent order. Values only identify group membership. Example: blood group (A, B, AB, O), marital status, product category.
- Ordinal: Categories with a meaningful order, but the gap between categories is not necessarily equal or measurable. Example: customer satisfaction rating (Poor, Average, Good, Excellent), education level, military rank.
- Interval: Numeric values with equal, meaningful gaps between them, but no true zero point, so ratios are not meaningful. Example: temperature in Celsius, where 0°C does not mean “no temperature,” and calendar year.
- Ratio: Numeric values with equal gaps and a true, meaningful zero point, so both differences and ratios are meaningful. Example: height, weight, income, and age, where 0 genuinely means “none” and 40 kg is twice as heavy as 20 kg.
| Scale | Identifies Categories | Has Order | Equal Intervals | True Zero | Example |
|---|---|---|---|---|---|
| Nominal | Yes | No | No | No | Gender, City, Product Type |
| Ordinal | Yes | Yes | No | No | Satisfaction Rating, Grade (A, B, C) |
| Interval | Yes | Yes | Yes | No | Temperature (°C), Calendar Year |
| Ratio | Yes | Yes | Yes | Yes | Height, Weight, Income, Age |
The practical payoff of the scale is that it tells you which statistics are defensible. Each level inherits everything permitted by the levels above it.
| Scale | Permissible summaries | Permissible comparisons |
|---|---|---|
| Nominal | Mode, frequency, proportion | Equality only (\(=\), \(\neq\)) |
| Ordinal | Adds median, percentiles, range of ranks | Adds order (\(<\), \(>\)) |
| Interval | Adds mean, standard deviation, correlation | Adds meaningful differences (\(-\)) |
| Ratio | Adds geometric mean, coefficient of variation | Adds meaningful ratios (\(\div\)) |
The most common measurement-level mistake is treating ordinal data as if it were interval data, for example averaging a 1-to-5 satisfaction rating and reporting “the average satisfaction was 3.4” as though the distance between “Poor” and “Average” is identical to the distance between “Good” and “Excellent.” Strictly, ordinal data supports the median and mode, not the arithmetic mean, though in practice many analysts do average Likert-scale data as a convenient approximation. Know that you are making that trade-off when you do it.
A quick test to identify the level of measurement: ask whether the values can be ordered (rules out nominal), whether the gaps between values are equal and measurable (rules out ordinal), and whether zero genuinely means “none of this quantity” (separates interval from ratio). Working through these three questions in order will correctly classify almost any variable you encounter.
A caution about the four-level scheme
Stevens’ typology is indispensable as a teaching device, but it has been seriously criticised, most influentially by Velleman and Wilkinson (1993), and a professional analyst should know why.
Their central objection is that scale type is not a fixed property of the data, it depends on what you are using the number for. A raffle ticket numbered 126 is nominal when it identifies a winner, ordinal when you ask who arrived earlier, and ratio when you count how many people came. The same recorded number changes scale with the question.
They further argue that using the scheme to forbid particular statistics is bad practice: real measurements rarely match a scale definition perfectly, demoting them to a lower level throws away information, and useful modern methods such as trimmed means sit between Stevens’ categories rather than inside one.
The sensible position is the one this book takes throughout. Use the four levels as a prompt to think about what your numbers mean before computing with them, not as a rulebook that decides your analysis for you.
2.7 5. Classification by Time Reference
The next question is when the observations were taken, and whether the same units are observed more than once. This determines what kind of change you are able to detect, and it is the criterion that most often decides whether a research design can answer the question being asked of it.
- Cross-sectional data: Many units observed at a single point in time. A survey of 500 customers this month. The order of the rows carries no meaning.
- Time series data: One unit observed repeatedly over many time periods. A single store’s monthly revenue for six years. Here the order of rows is the whole point.
- Pooled cross-sectional data: Several cross-sections taken at different times, but drawn from different units each time. Three independent customer surveys run in 2024, 2025 and 2026.
- Panel (longitudinal) data: The same units observed repeatedly over time. The same 500 customers surveyed each year for three years. A panel is balanced when every unit appears in every period and unbalanced when some units drop out.
Many units, one period
One unit, many periods
Different units each period
Same units, every period
Within cross-sectional designs, marketing research draws a further distinction:
- Single cross-sectional: One sample, measured once. The standard one-off survey.
- Multiple cross-sectional: Two or more samples, each measured once, at different times. Because the samples differ, comparisons are only valid at the aggregate level, not for any individual.
- Cohort analysis: A special multiple cross-sectional design in which the successive samples are drawn from the same defined group, such as everyone born in a given decade, allowing that group to be tracked as it ages even though different individuals are surveyed each round.
Example
A telecom operator wants to know whether a price change reduced churn.
A cross-sectional survey after the change tells you the churn rate now, but not whether it moved. Pooled cross-sections before and after tell you the aggregate rate moved, but not who moved. A panel that follows the same subscribers across both periods tells you which subscribers changed behaviour and lets you link that change to the price rise, which is the only one of the three that supports a credible causal story.
Panels buy that extra power at a real cost. Because the same people must keep participating, panels suffer attrition, with member dropout that can run to a fifth of the panel each year, and the members who remain are systematically different from those who leave. Long-serving panellists also change their behaviour simply because they know they are being observed. Treat a panel’s headline numbers as representative only after checking who has fallen out of it.
2.8 6. Classification by Structure
Classical statistics quietly assumes every dataset is a neat rectangle of rows and columns. Most of the data an organisation actually holds is not. Classifying by structure asks whether the data fits a predefined schema.
- Structured data: Organised in a predefined format, fitting neatly into rows and columns with a fixed schema. Spreadsheets, relational database tables, transaction records.
- Semi-structured data: No rigid tabular schema, but carrying tags or markers that identify its elements. JSON, XML, CSV with irregular fields, email (standard headers wrapped around free text).
- Unstructured data: No predefined format at all. Free text, images, audio, video, social media posts, call transcripts, sensor streams.
| Structured | Semi-structured | Unstructured | |
|---|---|---|---|
| Schema | Fixed, defined in advance | Flexible, self-describing | None |
| Typical formats | SQL tables, .xlsx, .csv
|
JSON, XML, log files | Text, images, audio, video |
| Usually stored in | Relational database, data warehouse | Document store, NoSQL database | Data lake, object storage |
| Ease of analysis | Directly analysable | Needs parsing | Needs extraction or a model |
| Share of enterprise data | Small minority | Growing | The large majority |
Industry estimates commonly put unstructured data at roughly 90% of everything an organisation generates, and growing several times faster than structured data. That figure is worth holding onto, because it explains why so much of modern analytics is really about converting unstructured data into structured data: sentiment scores extracted from reviews, categories assigned to support tickets, counts derived from images. Once converted, everything else in this book applies to it normally.
2.9 7. Classification by Organisation State
Data can also be classified by how far it has already been processed, and by how many variables you are looking at simultaneously.
By degree of summarisation:
- Raw (ungrouped) data: Individual observations exactly as recorded, in no particular order. The 847 individual transaction amounts from yesterday.
- Arrayed data: The same observations sorted into ascending or descending order, the first step towards making sense of them.
- Grouped data: Observations collected into classes with frequencies, such as “₹0 to ₹500: 214 transactions.” Grouping makes patterns visible but discards the individual values, which is why a mean computed from grouped data is an approximation of the true mean.
By number of variables under study:
- Univariate: One variable at a time. What is the average transaction value?
- Bivariate: Two variables and the relationship between them. Does transaction value differ by payment method?
- Multivariate: Three or more variables simultaneously. How do value, method, region, and customer tenure jointly relate?
This criterion is the bridge into the next three chapters. Chapter 2 covers classification and tabulation, Chapter 3 covers how to choose class intervals when grouping, and Chapter 4 builds the full frequency distribution. All three are, in this vocabulary, about moving data from the ungrouped state to the grouped state without losing more information than you have to.
2.10 8. Classification by Scale and Generation
The final lens asks how much data there is, how fast it arrives, and how it was produced. In February 2001 the analyst Doug Laney, then at Meta Group, characterised the emerging data management challenge along three dimensions, and this “three Vs” framing has become the standard definition of big data.
- Volume: The sheer quantity of data to be stored and processed.
- Velocity: The speed at which data arrives and must be acted upon.
- Variety: The range of incompatible formats, structures, and semantics that must be reconciled.
Two further Vs are now commonly added:
- Veracity: How trustworthy the data is, since scale magnifies rather than cures measurement error.
- Value: Whether the data actually supports a decision, the only dimension that ultimately justifies the other four.
Two related distinctions sit alongside the Vs:
- Human-generated vs machine-generated: A survey response is human-generated and arrives in small volumes; a stream of IoT sensor readings is machine-generated and arrives continuously. Machine-generated data is the main reason volume and velocity became problems at all.
- Batch vs streaming: Batch data is collected and processed in blocks, such as a nightly sales upload. Streaming data is processed as it arrives, such as fraud scoring on a payment as it happens. The distinction decides your entire technical architecture, and it is a property of the data flow, not of the variables themselves.
Big data is a description of engineering conditions, not a statistical virtue. A very large sample makes standard errors small, which means even trivial differences become statistically significant, while doing nothing whatsoever about bias. A biased sample of ten million is worse than an unbiased sample of a thousand, because it delivers a wrong answer with false confidence. Chapter 13 returns to this point in detail.
2.11 Putting It Together: One Dataset, Eight Lenses
Example
A retail chain holds a table of customer records. Each row is one customer, surveyed once, in March 2026.
| customer_id | city | satisfaction | orders_2025 | monthly_spend | signup_year | review_text |
|---|---|---|---|---|---|---|
| C-1041 | Chennai | Good | 12 | ₹8,450 | 2021 | “Delivery was quick, packaging poor” |
| C-1042 | Mumbai | Excellent | 3 | ₹1,200 | 2024 | “No complaints at all” |
| C-1043 | Delhi | Average | 27 | ₹19,300 | 2019 | “Prices went up this year” |
Running all eight criteria over each column produces the following:
| Column | Source | Nature | Countability | Scale | Time ref. | Structure | Organisation |
|---|---|---|---|---|---|---|---|
customer_id |
Internal secondary | Qualitative | — | Nominal | Cross-sectional | Structured | Ungrouped |
city |
Internal secondary | Qualitative (poly) | — | Nominal | Cross-sectional | Structured | Ungrouped |
satisfaction |
Primary (survey) | Qualitative | — | Ordinal | Cross-sectional | Structured | Ungrouped |
orders_2025 |
Internal secondary | Quantitative | Discrete | Ratio | Cross-sectional | Structured | Ungrouped |
monthly_spend |
Internal secondary | Quantitative | Continuous | Ratio | Cross-sectional | Structured | Ungrouped |
signup_year |
Internal secondary | Quantitative | Discrete | Interval | Cross-sectional | Structured | Ungrouped |
review_text |
Primary (survey) | Qualitative | — | Nominal | Cross-sectional | Unstructured | Ungrouped |
Three things in that table are worth pausing on.
signup_year is interval, not ratio, even though it is a plain number. Year 0 is a calendar convention, not an absence of time, so “2024 is 1.001 times 2022” is nonsense. Differences are fine: a customer who signed up in 2019 has been with the chain 5 years longer than one from 2024.
orders_2025 and monthly_spend are both ratio but differ in countability. Orders are counted and cannot be fractional; spend is measured and can be. This is exactly what decides bar chart versus histogram in Chapter 4.
review_text breaks the rectangle. Every other column is structured; this one is unstructured text sitting inside an otherwise tidy table. Analysing it means first converting it into something structured, such as a sentiment score or a set of topic flags, at which point it acquires a scale of its own.
Application
Declaring
satisfactionan ordered factor in R or an ordered categorical in pandas is the code-level expression of the measurement scale, and both languages then enforce it: the median is available through the category codes, while a plainmean()is refused outright. That refusal is not a limitation to work around, it is the ordinal scale of Section 4 being applied for you. Everystr()and every.dtypesis asking you to commit to a classification.
2.12 Data Quality: The Criterion Behind Every Criterion
Classifying a variable correctly tells you what you may do with it. It says nothing about whether the values themselves are any good. Data quality is conventionally assessed along six dimensions:
- Accuracy: Do the values correctly describe the real-world thing they claim to describe?
- Completeness: Is all the required data present, or are there gaps?
- Consistency: Do the same facts agree across systems and across records?
- Timeliness: Is the data current enough for the decision it will inform?
- Validity: Does each value conform to its permitted format, type, and range?
- Uniqueness: Is each real-world entity represented exactly once, with no duplicates?
These dimensions interact with the classifications in this chapter in ways that catch people out. A city column can be perfectly valid, every entry a real city, and still be inconsistent if “Bengaluru”, “Bangalore” and “BLR” appear as three separate categories for one place. A nominal variable with 200 distinct labels for 40 actual cities will quietly wreck every cross-tabulation built from it, and no amount of correct classification will detect the problem for you.
2.13 Why Data Types Matter for Analysis
Classification is not academic trivia, it is a gatekeeper that decides which tools you are allowed to use later in this book.
- Mode works for every level, nominal through ratio, since it only asks “what occurs most often.”
- Median requires at least ordinal data, since it depends on being able to rank values from smallest to largest.
- Mean, variance, and standard deviation require interval or ratio data, since they depend on the gaps between values being numerically meaningful.
- Ratios, percentage change, geometric mean, and coefficient of variation require ratio data specifically, since they depend on a true zero.
- Histograms require continuous data; bar charts are for categorical and discrete data.
- Time-series methods require a time-referenced design; running them on cross-sectional data produces confident nonsense.
Every measure of central tendency and dispersion covered later in this book inherits these restrictions directly from the classifications in this chapter.
Looking Ahead
With a working vocabulary for what data is, where it comes from, and the eight criteria by which it is classified, the next chapter turns to organizing raw, unsorted observations into arrays, and classifying and tabulating them so that patterns become visible.
Summary
| Concept | Description |
|---|---|
| Data Basics | |
| Data | Raw, unprocessed facts and figures collected about people, objects, or events |
| Information | Data that has been organized and given context so it becomes meaningful |
| Variable | A characteristic that can take different values across observations, such as age or income |
| Unit of Analysis | The entity each row of a dataset describes, such as a customer, transaction, or month |
| Metadata | Data about the data, such as collection date, units of measurement, and permitted values |
| 1. By Source | |
| Primary Data | Data collected firsthand by the analyst for a specific purpose, through surveys or experiments |
| Secondary Data | Data collected earlier by someone else for a different purpose, then reused for the current question |
| Internal Secondary Data | Secondary data already held inside the organisation, either ready to use or needing further processing |
| External Secondary Data | Secondary data from outside the organisation: published materials, computerised databases, or syndicated services |
| Syndicated Services | Commercially collected data sold to multiple clients, gathered from households or from institutions |
| Evaluating Secondary Data | Checking a secondary source for specifications, error, currency, objective, nature, and dependability |
| 2. By Nature | |
| Qualitative Data | Data describing a quality or category, such as gender or product type, that cannot be meaningfully averaged |
| Quantitative Data | Data describing a measurable quantity expressed as a number, such as income or number of orders |
| Dichotomous Variable | A categorical variable with exactly two categories, whose mean when coded 0 and 1 gives a proportion |
| Polytomous Variable | A categorical variable with three or more categories, such as state or payment method |
| 3. By Countability | |
| Discrete Data | Quantitative data that takes only specific, countable values with no meaningful values in between |
| Continuous Data | Quantitative data that can take any value within a range, limited only by measurement precision |
| 4. By Measurement Scale | |
| Nominal Scale | A measurement scale with categories but no order, such as blood group or marital status |
| Ordinal Scale | A measurement scale with a meaningful order but unequal or unmeasurable gaps, such as a satisfaction rating |
| Interval Scale | A measurement scale with equal, meaningful gaps but no true zero point, such as temperature or calendar year |
| Ratio Scale | A measurement scale with equal gaps and a true zero point, such as height, weight, or income |
| Permissible Statistics | Each scale permits the statistics of the scales below it plus one more: mode, then median, then mean, then ratios |
| Limits of the Four-Level Scheme | Scale type depends on how a number is used, so the four levels should guide thinking rather than forbid methods |
| 5. By Time Reference | |
| Cross-Sectional Data | Many units observed at a single point in time, where the order of rows carries no meaning |
| Time Series Data | One unit observed repeatedly across many time periods, where the order of rows is essential |
| Pooled Cross-Sections | Several cross-sections taken at different times from different units, comparable only in aggregate |
| Panel (Longitudinal) Data | The same units observed repeatedly over time, allowing individual change to be tracked |
| Cohort Analysis | A multiple cross-sectional design tracking a defined group over time, though sampling different individuals each round |
| Panel Attrition | The progressive dropout of panel members, which leaves the remainder systematically unrepresentative |
| 6. By Structure | |
| Structured Data | Data organised in a predefined schema of rows and columns, such as a database table or spreadsheet |
| Semi-Structured Data | Data with no rigid schema but carrying tags or markers, such as JSON, XML, or email |
| Unstructured Data | Data with no predefined format, such as free text, images, audio, or video |
| 7. By Organisation State | |
| Ungrouped Data | Individual observations exactly as recorded, before being sorted or collected into classes |
| Grouped Data | Observations collected into classes with frequencies, making patterns visible but discarding individual values |
| Univariate, Bivariate, Multivariate | Analysis of one variable, of the relationship between two, or of three or more simultaneously |
| 8. By Scale and Generation | |
| Big Data and the Vs | Volume, velocity, and variety, as defined by Laney in 2001, with veracity and value commonly added |
| Batch vs Streaming Data | Whether data is processed in blocks after collection or acted upon continuously as it arrives |
| Big Data Does Not Cure Bias | Large samples shrink standard errors but do nothing about bias, so a biased large sample misleads confidently |
| Data Quality | |
| Data Quality Dimensions | Accuracy, completeness, consistency, timeliness, validity, and uniqueness |