About Kappa and Classification Accuracy Analysis

“All models are wrong but some are useful.” — George E. P. Box

Introduction

Spatially explicit models have various applications in resource management, including the development of vegetation and wildlife habitat relationship predictive surfaces. Appropriate applications of these models are impossible without informed approaches to model development and accuracy assessment of the resulting data products. In the absence of incisive model development and error analysis, spatially explicit models may be applied in ways that confound, rather than illuminate, our understanding of vegetation land cover and wildlife habitat.

Accuracy assessment provides a means of gauging model performance, and end users who conduct accuracy assessments gain important information regarding model reliability and the suitability of the modeling process. Following Congalton and Green (1999), accuracy assessments are useful for: (1) ascertaining the quality of a predictive surface; (2) improving classification quality by identifying and correcting sources of error; (3) facilitating comparisons of algorithms, techniques, model developers and interpreters; and (4) determining the relevance of the data product in the decision-making process.

When accuracy assessments are conducted at all, generally only the overall accuracy metric is provided — yet the error matrix behind that single number supports a rich family of statistics that describe model performance far more completely. This page walks through that family: the error matrix itself, overall, user's and producer's accuracies, the Kappa statistic with its variance and tests, the per-class diagnostic metrics, the class-proportion weighting that turns an accuracy assessment into an unbiased area estimator, and the modern statistics that today's journals expect. The Classification Accuracy (Kappa) tool computes everything described here; this page explains what the numbers mean and where they come from.

The error matrix

The foundation of nearly everything on this page is the error matrix (also called a confusion matrix or contingency table). Each reference sample contributes to exactly one cell: the cell at row i, column j counts the sample points that were predicted to be class i and observed in the field to be class j. The diagonal (where i = j) represents cases where the predicted value agreed with the observed value; the off-diagonal cells contain the misclassifications, and their row and column describe exactly how each value was misclassified.

Observed (reference) class
12…k
Predicted
class
1n11n12…n1k
2n21n22…n2k
……………
knk1nk2…nkk
The general k-class error matrix. Row totals ni+ count the samples predicted as class i; column totals n+j count the samples observed as class j.
“Marginal,” in plain terms Statistics uses the word marginal in a way that has nothing to do with its everyday sense of “borderline” or “unimportant,” and the collision trips up many readers. The statistical term is purely about bookkeeping location: sum each row of a table like the one above and write the totals down the right edge; sum each column and write the totals along the bottom. Those totals sit, literally, in the margins of the table — so they are the marginal totals (or just “the marginals”). Each one answers a one-variable question by summing the other variable away: the row margins are what the classification claims of each class, ignoring the ground; the column margins are what the reference samples found, ignoring the classification. From there the jargon unwinds itself — marginal homogeneity just means the two sets of totals agree, and the chance-agreement term in Kappa is built from the marginal proportions. Wherever the word appears on these pages, substituting “the totals in the table's margins” will carry you through.

Overall, producer's and user's accuracy

The overall accuracy is simply the total number of correct classifications divided by the total number of sample points — the sum of the diagonal over the matrix total.

A useful companion to overall accuracy is the no-information rate: the accuracy you could achieve by ignoring the samples entirely and always guessing the commonest reference class. In the tutorial dataset the commonest class is Oak/Juniper, at 133 of the 362 reference samples, so a “classifier” that mindlessly labeled every sample Oak/Juniper would score 133 / 362 = 36.7% — no imagery, no predictors, no effort. That number is the floor any real classification must beat, and the report tests whether yours does (a one-sided exact binomial test of overall accuracy against the no-information rate, following the convention of R's caret package). Read it as a sanity floor rather than a high bar: failing it means the classification is extracting no usable information from its predictors, while passing it says nothing about whether the accuracy is good enough for your purposes. The rate itself deserves a glance too — the more imbalanced the landscape, the higher the floor sits, and the more flattering a raw overall accuracy can look, a theme that returns with force in the Kappa debate below.

Overall accuracy is often the only accuracy statistic reported with predictive landscape models, but the error matrix provides the means to calculate two further accuracies for each class — and these per-class accuracies are often more interesting to researchers than the overall figure, since a study species usually cares about particular classes rather than the classification as a whole. Indeed, the modern critics of summary statistics recommend the per-class accuracies as the centerpiece of the report (see the Kappa debate below). The distinction between the two matters:

The difference is best understood with an example. Consider this error matrix of 136 sample points over three vegetation classes:

Reference data
Deciduous
Forest
Coniferous
Forest
GrasslandTotal
Classification
data
Deciduous Forest6022486
Coniferous Forest230335
Grassland141015
Total635617136

The overall accuracy is (60 + 30 + 10) / 136 = 73.5%. But suppose we are interested specifically in deciduous forest. The reference data include 63 cases identified as deciduous forest, and 60 of them were correctly classified — a producer's accuracy of 95.2%, suggesting deciduous forest is captured very well. Notice, however, that 86 cases were classified as deciduous forest, so a “deciduous forest” label from the classification is correct only 60 / 86 = 69.8% of the time. Although the class can be classified with high accuracy from the producer's perspective, the actual reliability of the classification to the user may be much different — and both numbers deserve reporting.

The Kappa statistic

The Kappa statistic measures the agreement between two sets of categorizations of a dataset while correcting for chance agreement between the categories. Even a random classification would agree with the reference data some of the time, simply because both sides use the same classes in similar proportions; Kappa asks how much of the observed agreement exceeds that chance level. Following Congalton and Green (1999:50):

K^  =  n Σ nii − Σ ni+ n+in² − Σ ni+ n+i

with its variance estimated by the delta method — the θ1 through θ4 formulation of Fleiss, Cohen and Everitt (1969), where each θ term is computed directly from the matrix counts and marginals:

The Kappa estimator and its delta-method variance, with the theta-1 through theta-4 terms defined from the matrix counts and marginals
The Kappa estimator with its variance: the θ1–θ4 formulation (Fleiss, Cohen and Everitt 1969, as presented in Congalton and Green 1999).

The logic of the chance correction can be summarized as follows (after Agresti 1990:366–367). If Πo is the overall probability of correctly classifying a point (the overall accuracy), and Πe is the probability of agreement that would occur if the predicted and observed classifications were statistically independent — agreement purely by chance, computed from the products of the marginal proportions — then Πo − Πe is the excess accuracy beyond chance, and Kappa divides it by the excess that a perfect classification would have achieved:

K^  =  Πo − Πe1 − Πe

K^ typically ranges between 0 and 1, with values closest to 1 reflecting the highest agreement; negative values are possible but rare. For the example matrix above, K^ = 0.53 — the classification performs 53% better than chance alone would.

A caution Congalton and Green (1999:58–59) point out that some researchers object to Kappa for remote sensing accuracy because the chance-agreement term includes some agreement that is not purely due to chance, especially when the marginals are not fixed a priori — which is normally the case in remote sensing. We suspect it is rare that a researcher would decide a priori that X% of their landscape will be classified as Y. Even so, Kappa comes with powerful statistical properties that make it extremely useful — and a substantial modern debate, discussed below, about how it should be used.

Variance, significance tests and confidence intervals

Because K^ is asymptotically normally distributed, and because its variance can be estimated from the matrix (the formulation shown above), three important inferences follow:

One further test in the report deserves its own introduction: the McNemar–Bowker marginal-symmetry test (McNemar 1947; Bowker 1948), which asks a question none of the statistics above ask — not how much the classification errs, but whether its errors are directionally biased. Under the null hypothesis the mistakes are an even trade — the misclassifications mirror each other: for every pair of classes, about as many samples get misclassified from class A into class B as from B into A, and consequently the classification's class totals agree with the reference totals apart from sampling noise. (Note this is symmetry of counts, not of per-stand error rates — when one class is far commoner than the other, an even trade of counts actually means the rarer class is erring at the higher rate.) The statistic is built from the off-diagonal cells alone — each pair of classes contributes (nij − nji)² / (nij + nji) — on k(k − 1)/2 degrees of freedom, and for two classes it reduces exactly to McNemar's original paired test. A significant chi-square says at least one pair of classes is trading samples predominantly one way. (Strictly, one-way flows around a cycle of three or more classes could cancel in the totals, but that is a curiosity; in practice the asymmetry shows up as classified totals that differ from the reference.) In the tutorial dataset, for example, the classification claims 174 samples' worth of Mixed Conifer where the observer found 130, while shorting Pine/Oak 35 to 52 — and the test (chi-square 31.85 on 6 df, P = 0.00002) confirms that imbalance is real bias, not luck. This is the inferential cousin of the quantity disagreement component described below, and it is an early warning that the classification's raw acreages are biased — precisely the problem the error-adjusted area estimates correct. When the test is not significant, the classification's class proportions are statistically defensible as drawn, and most of whatever error exists is allocation — right amounts, wrong places.

Why would you want to know this about your classification method? Because the two outcomes describe different diseases, calling for different remedies. Symmetric errors mean the method is fuzzy but fair: it confuses two classes, but guesses evenhandedly when it does, so no acreage flows systematically from one class into another. The only real cure for fuzziness is better discrimination — better training data for the confusable pair, an additional predictor variable, imagery from a season when the classes separate — and adjusting thresholds will not help, because the amounts are already right. Asymmetric errors mean the method is slanted: something systematic — an over-generous decision threshold, training data that let a dominant class swamp its neighbors, a one-way spectral resemblance — is pushing acreage in a preferred direction, and everything downstream that consumes areas inherits the bias. Here recalibration genuinely helps, and the direction of the asymmetry tells you which way to lean: in the tutorial dataset the classification converts Pine/Oak into Mixed Conifer at five times the reverse rate, which smells like sparse pine-oak canopy reading as conifer, and says plainly that Mixed Conifer's signature needs reining in. Even before anything is fixed, a slanted classification errs predictably, which is itself usable: a manager who knows Pine/Oak is systematically swallowed knows the classified extent is a floor, not an estimate. One caution, in the spirit of Tukey's objection below: the P value only says the lean is real, not that it is large — with enough samples even a trivial lean turns significant — so read the size of the margin gaps and the quantity-disagreement share first, and the P value second.

Tukey's objection — why test at all? John Tukey famously observed: “All we know about the world teaches us that the effects of A and B are always different — in some decimal place — for any A and B. Thus asking ‘are the effects different?’ is foolish.” He was right — and the point applies here. Of course the classification differs from random chance at some decimal place, and of course two observers differ somewhere. So why test? Because the test does not really ask whether a difference exists; it asks whether your data are sufficient to establish its direction and rough size. A significant Z against chance is a sanity floor every assessment should clear — it says the sample is large enough to rule out the possibility that the agreement you see is the kind of thing chance produces, and with small samples that is not automatic. A non-significant comparison between two observers does not say the observers are identical; it says these data cannot tell them apart, so treating them interchangeably (pooling under the common Kappa, for example) is defensible for this analysis. This is why we recommend reading the confidence intervals first — they carry the magnitude and the uncertainty together — and treating P-values as summaries of what the data can and cannot distinguish, rather than as verdicts about the world. Tukey himself preferred the question “can we be confident of the direction of the difference?” — which is exactly the question those intervals answer.
Chebyshev's honest worst case The confidence intervals above are accurate if the statistics are normally distributed. If normality is in doubt — small samples, thin classes, or estimates near their limits — Chebyshev's Inequality still guarantees the interval covers at least 1 − 1/Z²α/2 of the probability: a nominal 95% interval is at worst a 74% interval. The tools report this worst-case coverage alongside every Kappa interval, so the reader is never left guessing; the optional bootstrap intervals (below) are the modern remedy when it matters.

Cohen's three independence assumptions — and what breaking them costs

Cohen (1960) laid down three independence assumptions for the statistic, and it is worth being precise about them, because the obvious worry — two observers rating the same points — is not a violation at all. Two judges categorizing the same sample of units is Cohen's design, not a breach of it. What he actually requires is independence of three other things, and each one, when violated, pushes the statistics in its own direction:

The same-points worry does surface in one legitimate place: not within either observer's analysis, but when their two kappas are compared. Kappas computed from overlapping samples are positively correlated, and the pairwise Z test ignores that covariance, overstating the variance of the difference — so the comparison runs conservative (see the Compare Accuracy Analyses page). Notice the tidy symmetry: violating unit independence within an analysis makes its tests overconfident, while sharing units between analyses makes their comparison underconfident — the same root cause, pushing in opposite directions depending on where the correlation lives. In both cases the point estimates survive; it is always the uncertainty statements that take the damage.

Per-class statistics: the confusion-matrix family

Several further metrics describe model performance for each class, derived by collapsing the error matrix to a 2×2 confusion matrix for that class — the four possible ways a sample may be classified and observed:

Observed XObserved not-X
Classified Xab
Classified not-Xcd
a = correct presences · b = false presences · c = missed presences · d = correct absences

The modern additions complete the family that R's caret package reports, and each earns its keep:

For detailed treatments of these concepts, the classic guide to the sensitivity/specificity family in ecological modeling remains Fielding and Bell (1997). The modern additions carry their own literature: the F-measure descends from information retrieval (Van Rijsbergen 1979); Sokolova and Lapalme (2009) give formal definitions for the family and — more usefully — work out which changes in the error matrix each measure is blind to; their observation that precision, recall and F1 never see the correct absences is exactly why the report pairs F1 with specificity, negative predictive power and balanced accuracy rather than letting it stand alone. Allouche, Tsoar and Kadmon (2006) examine how prevalence distorts kappa in species-distribution models and champion the prevalence-independent true skill statistic as the alternative (their TSS = sensitivity + specificity − 1 is a simple rescaling of the balanced accuracy reported here: TSS = 2 × balanced accuracy − 1); Brodersen et al. (2010) formalize balanced accuracy; and Kuhn (2008) documents the caret package whose report these statistics mirror. Weighted-average versions over all classes are computed by summing the per-class confusion matrices into one overall matrix; the per-class conditional Kappa (Congalton and Green 1999) applies the same chance correction classwise, with its own standard error.

Important: 2×2 presence/absence models Many researchers design a study to generate a simple two-category presence/absence habitat model. In that case the full error matrix is the confusion matrix, and the “overall” weighted-average statistics may not be what you want — with only two categories, the sensitivity of “presence” equals the specificity of “absence” and vice versa. If you are interested in how well the model predicts a species' habitat, refer to the statistics for the presence class; if in how well it predicts where the species will not be, refer to the absence class.

Designing the assessment

Classification considerations

The choice of classification system can greatly influence the measured accuracy. Congalton (1991) offers guidelines worth restating: (1) all areas to be classified should fall into one and only one category; (2) all areas should be classified — none left unclassified; (3) where possible use a hierarchical system, so categories can be collapsed into more general ones if the original set cannot meet your accuracy requirement; and (4) recognize that standard class breakpoints may be arbitrary — if your data cluster around a breakpoint (say, a canopy-closure class boundary), natural breaks may serve you better. (The Fisher-Jenks natural-breaks algorithm is built into ArcGIS Pro's own symbology tools — the “Natural Breaks (Jenks)” classification in any layer's symbology pane — so identifying sensible breakpoints takes nothing beyond the standard software.)

Sample point considerations

As a general rule of thumb, Congalton and Green recommend a minimum of 50 sample points per category, increasing to 75–100 per category with large numbers of categories (more than about 12). This may be an obvious point, but under no circumstances should you completely fail to sample any category — it is difficult to estimate the classification accuracy of a category you never checked. This suite of tools includes everything needed to plan and draw the design: the Estimate Sample Size tool computes the required totals by four methods, and the Random Point Generator and Select Random Records tools draw the design — stratified by the classification's categories, with spacing constraints, or from an existing sampling frame.

Adjusting for locational uncertainty

Locational uncertainty is a commonly unacknowledged source of error. A GIS assumes locations are perfectly precise and accurate, and the classification value at a sample point is normally extracted simply from the cell or polygon the point intersects. In the example below, the sample point is located on a Grassland cell (cell G), and by default the classification value of this point will be Grassland:

A sample point on a classified grid, landing on Grassland cell G among Pine, Oak and Mixed Conifer cells
The simple approach: the sample point takes the value of the single cell it intersects — here, Grassland.

But the coordinates of that point are probably not known exactly — if a GPS receiver determined the location, the point may well have landed on the wrong cell of a perfectly correct classification. If you suspect significant locational uncertainty in your samples, a more conservative approach is to extract the majority value in a circular area around each point — a correction suggested by Dr. Russell Congalton for the original extension, and available in the Classification Accuracy tool as the circle-radius option (ties go to the value at the point). Note that the same point is classified differently using the circular area instead of the cell it happens to intersect:

The same sample point with a circular neighborhood: summing the cell areas within the circle by class, Pine covers 6.791 hectares against Grassland's 5.517, so the point is classified as Pine
The circular-neighborhood approach: summing cell areas within the circle by class, Pine (6.791 ha) outweighs Grassland (5.517 ha), and the final classification is Pine.

As an aside: we strongly recommend reading past the “X-meter accuracy” blurb on the box your GPS came in. Most receivers define accuracy as a distance within which some percentage of fixes fall — and there is a very large difference between 50% of points within X meters (Circular Error Probable) and 95% of them.

Class proportions and error-adjusted areas

Everything above treats each sample point equally — which describes the sample. But the sample's class mix is rarely the landscape's class mix, and Card (1982) showed that weighting the matrix cells by each class's true landscape proportion πi restores unbiased landscape-level accuracies, with variance formulas for both simple and stratified random sampling designs. The original kappa_stats extension implemented exactly this as its “weight analysis based on classification value landscape proportion” option.

Olofsson et al. (2014) matured Card's weighting into the field's good-practice standard, with one decisive reframing: read in the other direction, the same machinery is an unbiased area estimator. The reference samples falling within each classified category tell you how that class's area truly divides among the ground classes — so the classification's own acreage figures can be corrected for misclassification, with standard errors.

A small example shows why this matters. Suppose a 10,000-ha classification shows 9,000 ha of forest and 1,000 ha of clearing, and the reference samples give a user's accuracy of 80% for classified clearing and 95% for classified forest — of the ground the classification calls clearing, 80% really is clearing; of the ground it calls forest, 95% really is forest. User's accuracy is the operative ingredient here, because the correction works from the classification's side: we stand on each classified category and ask what the reference samples found there. With only two classes, the 5% of classified forest that is not forest must be clearing, so the corrected clearing area is 1,000 × 0.80 + 9,000 × 0.05 = 1,250 ha — a full 25% more than the classification shows, because a small error rate applied to a large class moves more hectares than a large error rate applied to a small one. (Producer's accuracy — how much of the true clearing the classification managed to find — is what the correction produces, not what it consumes: with the corrected total in hand, the classification found 800 of the true 1,250 ha, a producer's accuracy of 64%.) No agreement statistic would ever reveal any of this; the area estimator is the tool that fixes the numbers you were going to quote from the classification. The Classification Accuracy tool measures the class proportions and the total classified area automatically from the classification dataset and reports each class's error-adjusted area in hectares with its standard error and confidence interval.

The modern statistics

The Kappa debate

A substantial modern literature argues against Kappa as a routine accuracy measure — most pointedly Pontius and Millones (2011, “Death to Kappa”) and Foody (2020). The complaints target the final summary statistic itself, not the analysis around it: Kappa measures agreement beyond chance rather than accuracy; its chance baseline is a poor model of how classifications actually err; and it is so sensitive to class prevalence that classifications with 95% overall accuracy can carry Kappa values anywhere from slightly below zero to 0.90 (Foody's own demonstration).

Rather than take that on faith, work through a small demonstration. Below are two error matrices of 200 samples each. In both classifications, every class is labeled with the same skill — and both have an overall accuracy of exactly 95%. The only thing that differs is how common the two classes are:

Obs. AObs. B
Pred. A955
Pred. B595
Classification 1 — balanced classes. Overall accuracy 95%; K^ = 0.90.
Obs. rareObs. common
Pred. rare55
Pred. common5185
Classification 2 — one rare class. Overall accuracy 95%; K^ = 0.47.

Same overall accuracy, yet Kappa falls from 0.90 to 0.47 — nearly halved — purely because one class became rare, which inflates the chance-agreement term. Now ask: which numbers actually tell you something you can act on? In Classification 2, the rare class's user's and producer's accuracies are both 5 / 10 = 50% — a coin flip. If that rare class is the wetland or the critical habitat your study cares about, the per-class table states the problem plainly, while the single Kappa of 0.47 only mutters something ambiguous: is this a mediocre classification, or a decent one applied to an imbalanced landscape? You cannot tell from Kappa alone — and that is the critics' entire point.

Hence what they recommend reporting instead: the full error matrix (which contains everything), overall accuracy with its interval, the per-class user's and producer's accuracies (which localize the problem), and the area-adjusted good-practice estimates (Olofsson et al. 2014; Stehman and Foody 2019 — which fix the acreage the errors distort), with McNemar's test as the recommended way to compare classifications evaluated on the same samples. Readers will notice that this is precisely the report the Classification Accuracy tool produces.

Kappa nonetheless keeps its adherents. It anchors the standard textbook treatment (Congalton and Green 1999) and decades of published accuracy statements that new work must be compared against — and even Foody grants it a legitimate remaining niche in assessing agreement among multiple interpreters, which is exactly the observer-comparison analysis these tools inherit from the original extension. For the prevalence paradox specifically, Gwet's AC1 is the robust alternative.

Our advice Report the matrix and the accuracies as the critics ask, and quote Kappa knowingly — for comparability with the existing literature and for observer-agreement questions — rather than as the sole headline figure. All models are wrong; some are useful. The same holds for their statistics.

Case study — Mexican Jay distribution surfaces, Pinaleños Mountains, Arizona

A brief case study, based on research conducted by Wynne (2003), illustrates how these statistics work together when neither model is statistically distinguishable from chance — a situation that may be familiar to anyone who models rare species. Two predictive habitat surfaces for the Mexican jay (Aphelocoma ultramarina) were compared: a classification-tree model built from a 1993–95 retrospective dataset, and a literature-based model built from published habitat accounts. Both models classify the same landscape into predicted habitat and predicted non-habitat, and both were assessed against the same field observation points:

The literature-based Mexican Jay habitat model: predicted habitat and non-habitat across the Pinalenos study area, with the field observation points overlaid
The literature-based model, with the field observation points overlaid.
The classification-tree (CART) Mexican Jay habitat model over the same study area and observation points
The classification-tree (CART) model over the same study area and observation points. Note how much more of the landscape it admits as habitat.

A statistical comparison suggested no significant difference between the two models (Z = 0.335, p = 0.369), and neither model's Kappa was significantly better than chance.

However, using a multi-criteria model selection process — highest overall accuracy, lowest misclassification rate, highest K^, highest sensitivity and specificity, highest predictive power, lowest commission and omission — the classification-tree model outperformed the literature-based model on nearly every criterion (overall accuracy 63.6% vs. 56.8%; K^ = 0.22 vs. 0.136; higher predictive power for both presence and absence). The case study provides a tractable method for identifying the model with the highest potential performance even when statistical significance is lacking — often the honest situation with limited field data.

The same analysis through the modern report

A friendly disclosure Twenty-some years on, we simply could not find a copy of the original CART raster — some datasets outlast their disk drives, and some do not. To demonstrate the modern statistics we reconstructed the surface from the published figure above: the figure was georeferenced, and the cells hidden underneath the sample-point markers (about 2.6% of the study area) were assigned the class of their nearest visible neighbor. The sample points in hand are also a slightly different vintage than the table Wynne published. So read the numbers below as a faithful demonstration of the report on nearly-original data, not as a re-publication of the 2003 findings — the two agree in character, and differ modestly in the decimals.

The two error matrices from the reconstructed analysis (44 field observation points; 11 observed presences, 33 observed absences):

Obs.
habitat
Obs.
nonhabitat
Pred. habitat716
Pred. nonhabitat417
CART model (reconstructed).
Obs.
habitat
Obs.
nonhabitat
Pred. habitat812
Pred. nonhabitat321
Literature-based model.
StatisticCART modelLiterature model
Overall accuracy (95% CI) 0.545 (0.389–0.696) 0.659 (0.501–0.795)
No-information rate; P(accuracy > NIR) 0.750; P = 0.999 0.750; P = 0.937
Cohen's Kappa (95% CI); Z; P 0.111 (−0.136–0.358); 0.88; 0.378 0.286 (0.024–0.547); 2.14; 0.032
Gwet's AC1 (SE) 0.136 (0.161)0.373 (0.149)
McNemar marginal symmetry χ² = 6.05; P = 0.014 χ² = 4.27; P = 0.039
Pontius quantity / allocation 0.273 / 0.1820.205 / 0.136
Habitat: user's / producer's accuracy 0.304 / 0.6360.400 / 0.727
Habitat F1; negative predictive power 0.412; 0.8100.516; 0.875
Error-adjusted habitat area (Olofsson) 10,030 ± 2,803 ha
(16,799 ha as drawn)
— (needs the model raster's proportions; that surface is likewise lost to time)
Models compared (pairwise) Z = 0.95, P = 0.341; common Kappa 0.193, P(equal) = 0.341
The modern report on the reconstructed Mexican Jay analysis; every number comes from one run per model of the Classification Accuracy (Kappa) tool. Following Wynne's own table convention, the stronger value in each comparable row is bold-faced (for McNemar, the weaker directional bias; for Pontius, the smaller disagreement). The no-information-rate row earns no bold: neither model clears that floor.

Reading the table with the tools of this page in hand, four things stand out that the original overall-accuracy-and-Kappa summary could not show. First, neither model clears the no-information rate: with 33 of 44 observations being absences, always guessing “nonhabitat” scores 75%, a floor neither model's accuracy approaches — the rare-species prevalence problem in its purest form, and the reason the per-class rows matter more than the overall ones here. Second, both models fail the marginal-symmetry test in the same direction: each predicts roughly twice as much habitat as the observers found (23 and 20 predicted presences against 11 observed), a directional bias — though arguably the forgiving direction for a species of conservation concern, since habitat over-drawn invites a survey while habitat omitted invites a bulldozer. Third, the error-adjusted area quantifies that generosity for the CART surface: the reconstruction draws 16,799 ha of habitat, while the samples support an estimate of only about 10,030 ± 2,803 ha. And fourth, the punchline the original study reached by other means still stands: although the literature model's Kappa is individually distinguishable from chance here and the CART model's is not, the pairwise test cannot tell the two models apart (P = 0.341) — forty-four samples simply lack the power — so a defensible choice between them still rests on the multi-criteria comparison, exactly as Wynne concluded. The alert reader will notice that the bold now sweeps the literature model's column, where Wynne's 2003 table favored the CART model on nearly every criterion. We would not read too much into the reversal — a reconstructed surface, a different sample vintage, and only 44 points make the ranking itself delicate, which is precisely why the significance test declines to pick a winner — but it is a fair illustration of how the multi-criteria sweep works: line up every statistic, bold the better column, and let the preponderance speak, while holding the verdict as lightly as the sample size deserves.

References