Kappa Analysis Tutorial

A complete worked accuracy assessment with the Classification Accuracy (Kappa), Estimate Sample Size and Compare Accuracy Analyses tools · theory in About Kappa analysis
Sample data This tutorial works through the same Pinaleños Mountains dataset used in the classic extension's manual, packaged as a tutorial geodatabase (Kappa_Tutorial_Data.gdb): a supervised vegetation classification (as a raster and as polygons) and 362 field samples surveyed independently by two observers. Download the tutorial data here (0.3 MB) — unzip it anywhere convenient, add the geodatabase's datasets to an ArcGIS Pro map, and you can follow along click for click.

This tutorial follows one complete accuracy assessment from start to finish: planning the sample size, drawing the sampling design, running the assessment once per field observer, reading each block of the report, comparing the two analyses — first by hand with the Compare Accuracy Analyses tool, then automatically in a single run — and confirming the results independently in R. The steps form a natural workflow — but you are not chained to it. Every tool accepts generic inputs, so you can enter anywhere: bring your own samples, type a matrix's worth of values, or compare kappas straight from the literature.

The data

The tutorial geodatabase holds five datasets. The classification is a supervised vegetation classification of a portion of the Pinaleños Mountains, Arizona, with four classes: Mixed Conifer forest (35.7% of the classified area), Oak/Juniper woodland (43.0%), Bare Ground/Grass (13.2%) and Pine/Oak woodland (8.1%). It comes in two interchangeable forms — Classified_Raster, whose attribute table carries the class names in a field named S_VALUE, and Classified_Polygons, the same classification as polygons with the names in VEG_LABEL — so you can run the assessment against whichever form your own projects use.

The field samples are 362 surveyed locations, each visited independently by two observers who recorded the vegetation class they found in a field named COARS_VEG. They appear three ways: Observer_1 and Observer_2 hold each survey separately, and Observers_1_and_2 stacks both surveys in one feature class (724 points) with an Observer field telling them apart — the shortcut dataset for Step 8. (Why the samples were collected the way they were is a story for Step 1.)

Step 1 — How many samples would we need?

In a perfect world, we would start the research by generating random sample locations across the landscape. The randomness allows us to apply the knowledge we gain to the larger ecological system, rather than settling for a single-shot analysis that only tells us about our sample points. Ideally we would be able to reach and survey all our sample points, and we would save costs and resources by sampling only the minimum number necessary to achieve the confidence level we want from our statistics.

In this perfect world, we would use our Estimate Sample Size tool to size our sampling array before we began our field surveys. For example, suppose we had initially classified our landscape into 4 important categories, and we therefore have an estimate of the proportion of the landscape that lies in each category. Suppose also that we want our field survey effort to tell us the classification accuracy with 95% confidence and a 10% level of precision — meaning that we are willing to accept an accuracy estimate that lands within 10 percentage points of the true accuracy, and we want a survey design that would deliver an estimate that close in 95 out of every 100 repetitions. Given these initial conditions, we can use the Estimate Sample Size tool to tell us the number of sample points necessary, given that we know the proportion covered by the largest category (Oak/Juniper, at 43%):

ClassHectaresProportion
Oak/Juniper18,32143.0%
Mixed Conifer15,22435.7%
Bare Ground/Grass5,60613.2%
Pine/Oak3,4668.1%
Total42,616100%
The Estimate Sample Size dialog with the Largest class proportion known method, 4 map classes, 95 percent confidence, 10 percent precision, and a largest class proportion of 43
Estimation method: Largest class proportion known, with 4 classes, 95% confidence, 10% precision, and the 43% Oak/Juniper proportion.

Based on this, we see that we would need 153 randomly-distributed points (and stratified within categories, so that 43% of the random points fall within that Oak/Juniper category). You can find this value in the Tool Results window:

The Estimate Sample Size messages window showing the recommended sample size of 153 highlighted
The tool's messages: the multinomial calculation, the recommended sample size of 153, and the 50-per-class reminder.

Refer to the Estimate Sample Size online help page for details on how this tool can be run given different analytical goals. If we knew nothing about the relative proportions covered by the landscape categories, for example, we would need 156 randomly-distributed points. And if our goal were not to estimate our classification accuracy, but rather to estimate the true areas covered by each category on the landscape, then we would set up the Estimate Sample Size tool to target a standard error of overall accuracy (the Olofsson et al. 2014 method): pinning overall accuracy to within ±1 percentage point of standard error takes roughly 774 points if we are optimistic about our conjectured user's accuracies (85–95%), and over 2,200 points if we guess pessimistically near 50%. That is an honest planning lesson worth absorbing: inaccurate classifications need far more samples to pin down.

In the real world, when we originally did this analysis, we were limited by rough topography and simply unable to survey randomly, and consequently we applied a systematic approach by sampling at regular intervals along navigable trails. We knew our study design was weakened by this lack of randomness, so we sampled more points than we would have originally estimated to be necessary. Our 362 samples comfortably exceeded the multinomial requirements and give about 90 samples per class on average — well above Congalton and Green's 50-per-class rule of thumb. Had we needed to draw the design, the Random Point Generator (stratified by the classification's categories) or Select Random Records would be the next stop, and the recommended sample size flows into them directly in ModelBuilder:

A ModelBuilder model chaining Estimate Sample Size, its derived Recommended sample size output, the Classified Raster, and the Random Point Generator into a Random Points output
Estimate Sample Size's derived Recommended sample size output feeding the Random Point Generator in ModelBuilder — no retyping between planning and drawing the design.

Step 2 — Drawing the design with the Random Point Generator

We did not actually generate random points for this study — the topography vetoed that, as it may for yours. But in the interests of showing the complete analytical flow path, here is how we would draw the 153 random points, stratified by the proportion of the landscape in each category.

Open the Random Point Generator (Geometric Tools group). In the tool window:

The filled dialog should look something like the following:

The Random Point Generator dialog set to stratify by raster category with counts proportional to stratum area and 153 samples
The Random Point Generator, set to stratify by the classified raster's categories with counts allocated proportional to stratum area.

And the result is the sampling distribution we would have wished for — 153 random points, spread so that each vegetation category receives its proportional share:

A map of the classified Pinalenos landscape with 153 random sample points distributed proportionally across the four vegetation categories
Our dream sampling design: 153 stratified random points across the four vegetation categories.

The tool has considerably more depth than we needed here — minimum spacing between points, boundary offsets, other stratification and allocation recipes — all covered on the Random Point Generator help page.

Step 3 — Go get the data

Probably the most fun part of the whole analysis, for most of us: actually going out and surveying the sample points. Time to go get the data. There is no dialog to fill in for this step and no button to click — just boots, datasheets and a GPS unit. For this study it meant working through some genuinely rugged Pinaleños country, twice over (once per observer), recording the vegetation class actually found at each of 362 locations. We will not pretend that is a small investment of time and effort — it is usually the largest line in the whole project budget. But let's be honest with ourselves: days in the field like these are why most of us went into ecological analysis in the first place. The statistics in the steps that follow are simply how we make sure all that boot leather ends up telling us something durable about the landscape.

For the tools downstream, the field protocol reduces to one essential: at each sample point, record the class you actually find in a field of its own (ours is COARS_VEG) — independently of what the classification claims. Everything else you care to record while you are standing there is a bonus.

Step 4 — Running the assessment on Observer 1

Back at the computer: remember that our actual samples are not the random points of Step 2 but the 362 systematically-surveyed locations, and the first survey to assess is Observer 1's. Open Classification Accuracy (Kappa) (Ecological Analysis group). In the tool window:

The filled dialog should look something like the following (the dialog is tall, so it is shown here in two pieces — the top half on the left, the lower half on the right):

The Classification Accuracy (Kappa) dialog in two pieces: the top half with Observer_1, Classified_Raster, S_VALUE and COARS_VEG entered, and the lower half with the map-class-proportion weighting checked and Stratified random selected
The Classification Accuracy (Kappa) dialog, filled for Observer 1. Everything under Options and Advanced, and every unhighlighted entry, stays at its default.

Step 5 — Reading the report

The tool prints the full report in the messages (and to a text file if you asked for one). Walking through Observer 1's blocks in order:

The error matrix

Observed (reference)
Bare Ground/
Grass
Mixed
Conifer
Oak/
Juniper
Pine/OakTotal
Predicted
(classification)
Bare Ground/Grass16312334
Mixed Conifer81152526174
Oak/Juniper2277911119
Pine/Oak15171235
Total4713013352362

Read the off-diagonal cells before any statistic, because they tell several different stories at once. The largest single error cell is 26: samples where the crew stood in Pine/Oak but the classification said Mixed Conifer. Only 52 Pine/Oak samples were observed in total, so that one cell has already swallowed half the class — the main reason Pine/Oak's producer's accuracy will look so poor below. Summing each pair of classes in both directions tells another story: Oak/Juniper and Bare Ground/Grass trade 34 samples between them (12 misclassified in one direction, 22 in the other), the largest two-way confusion on the matrix. And Oak/Juniper trades 32 with Mixed Conifer (25 + 7), which is worth a pause, because those two types generally do not even neighbor each other on the mountain: the vegetation follows the elevation gradient, with Oak/Juniper at the lower elevations, Pine/Oak in the middle and Mixed Conifer up high (Bare Ground/Grass can appear anywhere), so the classification is confusing classes from opposite ends of the gradient, not blurring a shared boundary. Three observations from one table, and not one of them visible in any overall statistic.

Overall statistics

Per-class statistics

The per-class table is where management decisions live, so let's not rush past it. The key columns for Observer 1:

Classn predictedn observed User's
accuracy
Producer's
accuracy
F1Balanced
accuracy
Bare Ground/Grass3447 0.47060.34040.39510.6416
Mixed Conifer174130 0.66090.88460.75660.8152
Oak/Juniper119133 0.66390.59400.62700.7097
Pine/Oak3552 0.34290.23080.27590.5783
Per-class statistics, Observer 1. The report's table also carries conditional Kappa, sensitivity, specificity, the predictive powers and prevalence for every class.

Take one row and read it end to end. Pine/Oak was predicted at 35 samples but observed at 52, and its diagonal cell in the error matrix holds 12. User's accuracy reads the classification's own claims: of the 35 samples it calls Pine/Oak, 12 really are — 12/35 = 0.3429, so a manager standing on ground classified as Pine/Oak has roughly a one-in-three chance of actually being in it. Producer's accuracy reads the ground: of the 52 Pine/Oak stands the observer found, the classification caught 12 — 12/52 = 0.2308, so more than three-quarters of the real Pine/Oak is missing from the classified dataset. If Pine/Oak matters to your study species, this single row matters more than any overall statistic on the page.

Now contrast Mixed Conifer, where the two accuracies split in the opposite way: producer's accuracy is excellent (115/130 = 0.8846 — when you are standing in Mixed Conifer, the classification almost always knows it) while user's accuracy is mediocre (115/174 = 0.6609 — a third of what it calls Mixed Conifer is really something else, much of it that missing Pine/Oak). Reading the two accuracies as a pair tells you not just how good a class is, but which direction it fails in: Mixed Conifer is over-claimed, Pine/Oak is swallowed.

The remaining columns condense that pair into single scores. F1 is the harmonic mean of user's and producer's accuracy — a class only scores well if both are good, and the harmonic mean punishes imbalance (Mixed Conifer's 0.88/0.66 pair earns 0.76, not the 0.77 a plain average would give; weaker pairs are punished harder). Balanced accuracy averages how well the class's members are found (sensitivity, which is producer's accuracy by another name) with how well non-members are excluded (specificity); 0.5 is coin-flip performance, so Pine/Oak's 0.5783 sits uncomfortably close to guessing.

Below the table the report prints two summary lines that are easy to skim past but worth understanding:

Area-weighted estimates and error-adjusted areas

Because the analysis was weighted by the class proportions, the report finishes with the Olofsson block. The tool measured the total classified area automatically — 42,615.8 ha of analyzed classes — and corrected each class's acreage for the misclassification the matrix revealed:

ClassClassification saysError-adjusted estimate
Bare Ground/Grass13.2% (5,606 ha)16.0% ± 2.0% (6,824 ± 857 ha, or 5,967 to 7,681 ha)
Mixed Conifer35.7% (15,224 ha)28.5% ± 1.8% (12,129 ± 760 ha, or 11,369 to 12,889 ha)
Oak/Juniper43.0% (18,321 ha)42.3% ± 2.5% (18,012 ± 1,051 ha, or 16,961 to 19,063 ha)
Pine/Oak8.1% (3,466 ha)13.3% ± 1.8% (5,651 ± 752 ha, or 4,899 to 6,403 ha)
Error-adjusted area estimates (Olofsson et al. 2014), Observer 1, shown ± one standard error. Note the percentage SEs are in points of the total classified area, so ±2.0% of 42,616 ha is the ±857 ha beside it. The report also prints full confidence intervals (±1.96 standard errors at the 95% level).

This is the block worth pausing over. The classification under-represents Pine/Oak by roughly forty percent of its true extent — real Pine/Oak is being absorbed into the Mixed Conifer class — while Mixed Conifer is over-claimed by about a fifth. Kappa told us agreement was moderate; the area estimates tell us specifically which hectares to distrust, with error bars suitable for a manuscript.

Step 6 — The second observer

Run the tool a second time with Sample points or table set to Observer_2 and every other entry unchanged. The report has the same structure; the headline numbers land close to Observer 1's:

Observer 1Observer 2
Overall accuracy0.6133 (222 of 362)0.5994 (217 of 362)
Cohen's Kappa0.43170.4131
Kappa variance0.0011840.001134
Gwet's AC10.50090.4825
The two independent runs, side by side. Note the kappa and its variance for each — those two numbers per analysis are all the next step needs.

The same story holds class by class — Mixed Conifer strong from the producer's side, Pine/Oak weak from both — which is itself informative: the failings look like properties of the classified dataset, not of whoever did the surveys.

Step 7 — Comparing the analyses

Are the two observers statistically distinguishable? Open Compare Accuracy Analyses and type one row per analysis into the value table, straight from the two reports: Observer 1 with kappa 0.4317 and variance 0.001184, Observer 2 with kappa 0.4131 and variance 0.001134. The tool prints:

The filled tool and its results should look something like the following:

The Compare Accuracy Analyses dialog with two rows, Observer 1 with kappa 0.4317 and variance 0.001184 and Observer 2 with kappa 0.4131 and variance 0.001134, beside its results window with the pairwise and common-kappa tests highlighted
The Compare Accuracy Analyses dialog with the two typed rows, and its results: each kappa's Z test against chance, the pairwise comparison, and the Fleiss common kappa.

The two observers are statistically indistinguishable — a reassuring result, and precisely the multiple-interpreter agreement question that even Kappa's critics agree the statistic answers well (see the Kappa debate). Notice what this tool asked for: just a kappa and a variance per analysis. Nothing requires the analyses to share sample points, sample sizes, software, or even a decade — two crews surveying entirely different plots compare just as naturally, and you can test your model against a kappa published in the literature by typing two rows.

One honest nuance runs the other way. The Z and common-kappa tests assume the analyses are independent samples — which different plots satisfy exactly. Our two observers stood at the same 362 points, so their kappas are positively correlated and the test runs slightly conservative: real differences must be a little larger than usual to register. A significant result can therefore be trusted; a borderline non-significant one deserves a second look with McNemar's paired test (Foody 2004), the sharp tool for same-sample comparisons. Here, at P = 0.700, the observers agree by either reading.

Step 8 — The one-run shortcut

Because both surveys live in one feature class, we could have done Steps 4 through 7 in a single run — the samples need not share locations, only a table and a field that tells the analyses apart. Point the tool at Observers_1_and_2 and set the Analysis field (under Options) to Observer. The tool splits the samples by that field's values, produces one complete report block per observer (identical, number for number, to the separate runs above), and finishes with the same comparison block: Z = 0.386, P = 0.700, common kappa 0.4222. The filled dialog (tall, so shown in three pieces) and the tail of its results:

The Classification Accuracy (Kappa) dialog in three pieces, with Observers_1_and_2 as the sample points, Classified_Raster and S_VALUE for the classification, COARS_VEG for the observed field, the Observer analysis field selected under Options, and the map-class-proportion weighting checked
The one-run setup: identical to Step 4 except the sample points are Observers_1_and_2 and the Analysis field (under Options) is set to Observer.
The results window of the combined run, ending with the highlighted Comparisons between analyses block: 1 vs 2 Z = 0.386, P = 0.69981, and the Fleiss common kappa of 0.4222
The tail of the combined run's report: after each observer's complete block, the automatic comparison — the same pairwise Z, P and common kappa as Step 7, to the digit.

Use whichever form fits your bookkeeping. Separate runs keep each model or observer in its own dataset and report — and Compare Accuracy Analyses will reconcile them later, even from typed values. The analysis field suits samples that already live in one table: per observer, per model, per imagery date — one run, comparisons included, and a single R snippet that reproduces everything.

Step 9 — Confirming the results in R

Every report ends with a ready-to-run R snippet carrying your own error matrix and strata weights. Paste it into R and it reproduces the analysis with the reference packages — psych for Kappa and its variance, caret for the per-class suite, diffeR for the Pontius disagreements, irrCAC for AC1, and mapaccuracy for the Olofsson areas. Those five packages are the only setup R needs, installed once:

install.packages(c("psych", "caret", "diffeR", "irrCAC", "mapaccuracy"))

With an analysis field, the snippet carries every observer's matrix and finishes with the same pairwise Z and common-Kappa tests. Running it for this dataset confirms each number above to the printed digit — we encourage the skeptical reader to try; skepticism is, after all, what this whole analysis is for.

Where to go from here