Footnotes — What the Cited Authors Actually Said

Supporting statements for the citations in About Kappa analysis, the Classification Accuracy (Kappa) tool page and the tutorial
Why this page exists Every substantive claim on these pages cites a source, and this page shows what those sources actually say — verbatim, with page numbers — so a skeptical reader can check our claims against the original authors' words without a library trip. Assembling it is good discipline for us, too: several passages on these pages were corrected during this verification, and we would rather show the receipts than ask for trust. The small ¶ marks in the body text jump here. Entries appear as we complete the verification; papers not yet verified are listed at the bottom.

Agresti (1990)

Agresti, A. 1990. Categorical Data Analysis. John Wiley and Sons, New York. (Kappa: pp. 364–367.)

We cite it for: the numbered logic behind the Kappa derivation on the About page (“after Agresti 1990:366–367”), including the chance-agreement term built from the marginals and the Fleiss et al. variance.

“Kappa equals 0 when the agreement equals that expected by chance, and it equals 1.0 when there is perfect agreement. The stronger the agreement, the higher the value, for a given pair of marginal distributions.” — p. 366
“It is rarely plausible that agreement is no better than expected by chance. Thus, rather than testing H0: κ = 0, it is more important to estimate strength of agreement, by constructing a confidence interval for κ.” — p. 367 — direct support for this site's read-the-intervals-first advice
“We distinguish between measuring agreement and measuring association, because there can be strong association without strong agreement.” — pp. 365–366

Allouche, Tsoar and Kadmon (2006)

Allouche, O., A. Tsoar, and R. Kadmon. 2006. Assessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (TSS). Journal of Applied Ecology 43:1223–1232. doi.org/10.1111/j.1365-2664.2006.01214.x

We cite it for: the demonstration that prevalence distorts kappa in species-distribution models, and the true skill statistic (TSS) as the prevalence-independent alternative (About — the per-class reading list; TSS = 2 × balanced accuracy − 1).

“The theoretical analysis shows that kappa responds in a unimodal fashion to variation in prevalence and that the level of prevalence that maximizes kappa depends on the ratio between sensitivity… and specificity… In contrast, TSS is independent of prevalence.” — Summary point 3, p. 1223
“We therefore recommend the TSS as a simple and intuitive measure for the performance of species distribution models when predictions are expressed as presence–absence maps.” — Summary point 5, p. 1223

Bowker (1948)

Bowker, A. H. 1948. A test for symmetry in contingency tables. Journal of the American Statistical Association 43:572–574. doi.org/10.1080/01621459.1948.10483284

We cite it for: the k-class marginal-symmetry test printed in every report of the Classification Accuracy (Kappa) tool (About — the McNemar–Bowker passage), including the off-diagonal-only statistic, the k(k−1)/2 degrees of freedom, and the reduction to McNemar's test for two classes.

“It is desired to test the hypothesis that the above table is symmetric (i.e., Pij=Pji). If this hypothesis is true, then Pk.=P.k; that is, the marginal distribution of individuals is the same for classifications A and B.” — p. 572
“[The statistic] has for large N the Chi-square distribution with m(m−1)/2 degrees of freedom… Note that the value of χ² is computed from non-diagonal frequencies, which is reasonable since diagonal frequencies should not afford evidence tending to disprove the hypothesis of symmetry.” — p. 573
“Formula (2) was obtained by McNemar [1], who lists applications in psychological testing.” — p. 573, on the 2×2 case

A charming aside: Bowker developed the test for a study of plate patterns on the left and right sides of stickleback fish — symmetry in the most literal biological sense.

Brodersen, Ong, Stephan and Buhmann (2010)

Brodersen, K. H., C. S. Ong, K. E. Stephan, and J. M. Buhmann. 2010. The balanced accuracy and its posterior distribution. Pages 3121–3124 in Proceedings of the 20th International Conference on Pattern Recognition. doi.org/10.1109/ICPR.2010.764

We cite it for: formalizing balanced accuracy as the corrective for accuracy's optimism on imbalanced data (About — the per-class reading list; the balanced-accuracy column of the per-class table in the Classification Accuracy (Kappa) tool's report).

“[Averaging accuracies across cross-validation folds] leads to an optimistic estimate when a biased classifier is tested on an imbalanced dataset. We show that both problems can be overcome by replacing the conventional point estimate of accuracy by an estimate of the posterior distribution of the balanced accuracy.” — Abstract, p. 3121

Card (1982)

Card, D. H. 1982. Using known map category marginal frequencies to improve estimates of thematic map accuracy. Photogrammetric Engineering and Remote Sensing 48:431–439.

We cite it for: the original derivation of weighting the error matrix by the classification's class proportions, with variances for both the simple-random and stratified-random designs — the two choices in the Classification Accuracy (Kappa) tool's Sampling design parameter (About — Class proportions and error-adjusted areas).

“Frequently, one has knowledge of the true map category marginal proportions, that is, the relative areas of each map category. These map category proportions can be used to improve estimates of ‘proportion-correct’ for each map category. This paper derives these improved estimates with their asymptotic variances for two common sampling designs: simple random sampling of single points and sampling stratified by map category.” — Abstract, p. 431
“As pointed out by Switzer (1969), a map without an accompanying statement of error is like a point estimate with no variance stated.” — p. 431

Cohen (1960)

Cohen, J. 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20:37–46.

We cite it for: Kappa itself — this is the founding paper, defining the statistic every page in this suite is built around, along with its hypothesis tests and confidence limits.

“The most primitive approach has been to simply count up the proportion of cases in which the judges agreed, po, and let the issue rest there… It takes relatively little in the way of sophistication to appreciate the inadequacy of this solution. A certain amount of agreement is to be expected by chance…” — p. 38 — the chance-correction rationale behind everything on these pages
“The purpose of this article is to present a coefficient to measure the degree of agreement in nominal scales, and to provide means of testing hypotheses and setting confidence limits for this coefficient.” — p. 39
“These have in common the following conditions, which may be taken as assumptions of the coefficient of agreement to be proposed: 1. The units are independent. 2. The categories of the nominal scale are independent, mutually exclusive, and exhaustive. 3. The judges operate independently.” — p. 38 — the three independence assumptions unpacked in the About page's assumptions passage
“…for any problem in nominal scale agreement between two judges, there are only two relevant quantities: po = the proportion of units in which the judges agreed… pe = the proportion of units for which agreement is expected by chance.” — p. 39 — the two ingredients of the Πo/Πe formula on the About page

Cohen's setting was psychiatric diagnosis — two clinicians sorting test protocols into categories, the judgment made by what he calls, quoting Stevens, a “two-legged meter.” Swap the clinicians for a classified raster and a field crew, and it is exactly our problem: the same statistic crossed from psychology into remote sensing and has never left.

Cohen (1968)

Cohen, J. 1968. Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin 70:213–220.

We cite it for: the weighted kappa the Classification Accuracy (Kappa) tool offers for ordinal classes (habitat suitability ranks, cover classes), where near-misses should count more than gross errors.

“A generalization to weighted kappa (KW) is presented. The KW provides for the incorporation of ratio-scaled degrees of disagreement (or agreement) to each of the cells… such that disagreements of varying gravity (or agreements of varying degree) are weighted accordingly. Although providing for partial credit, KW is fully chance corrected.” — Abstract, p. 213

Congalton (1991)

Congalton, R. G. 1991. A review of assessing the accuracy of classifications of remotely sensed data. Remote Sensing of Environment 37:35–46. doi.org/10.1016/0034-4257(91)90048-B

We cite it for: the assessment-design guidelines restated in the Designing-the-assessment section — classification system, sampling scheme, sample size — all organized around the error matrix.

“This paper reviews the necessary considerations and available techniques for assessing the accuracy of remotely sensed data. Included in this review are the classification system, the sampling scheme, the sample size, spatial autocorrelation, and the assessment techniques. All analysis is based on the use of an error matrix or contingency table.” — Abstract, p. 35

Congalton and Mead (1983)

Congalton, R. G., and R. A. Mead. 1983. A quantitative method to test for consistency and correctness in photointerpretation. Photogrammetric Engineering and Remote Sensing 49:69–74.

We cite it for: the classic workflow the Compare Accuracy Analyses tool implements — using kappa's variance to test whether interpreters, methods, or imagery conditions differ significantly.

“A method has been developed to quantitatively test the degree of similarity between photointerpreters and/or photointerpretation variables, such as film/filter type, season, and scale… This technique allows the comparison of individual interpreters, a test for photointerpreter consistency, and the comparison of photointerpretation variables.” — Abstract, p. 69
“The procedure proposed in this paper can test for significant differences between photointerpreters, test the consistency of the same interpreter over time, or test for significant differences between photointerpretation variables such as film type, season, and scale.” — p. 69

A historical aside: their analysis ran on a computer program named simply KAPPA — forty-odd years later, the same comparisons run in a dialog box.

Fielding and Bell (1997)

Fielding, A. H., and J. F. Bell. 1997. A review of methods for the assessment of prediction errors in conservation presence/absence models. Environmental Conservation 24:38–49. doi.org/10.1017/S0376892997000088

We cite it for: the classic guide to the sensitivity/specificity family for ecological presence/absence models — the framing behind the per-class 2×2 suite and the two-class guidance in the Classification Accuracy (Kappa) tool's report.

“The simplest, and most widely used, measure of prediction accuracy is the number of correctly classified cases. There are other measures of prediction success that may be more appropriate.” — Summary, p. 38
“A range of techniques for measuring error in presence/absence models, including some that are seldom used by ecologists (e.g. ROC plots and cost matrices), are described… Thirteen recommendations are made to enable the objective selection of an error assessment technique for ecological presence/absence models.” — Summary, p. 38

Fleiss (1981)

Fleiss, J. L. 1981. Statistical Methods for Rates and Proportions, 2nd ed. John Wiley and Sons, New York. Chapter 13, The Measurement of Interrater Agreement, pp. 212–236.

We cite it for: the variance-weighted common kappa and its chi-square test that several analyses share one underlying kappa, printed by the Classification Accuracy (Kappa) tool whenever an analysis field splits the samples and by the Compare Accuracy Analyses tool from typed kappa and variance pairs.

“Suppose one wishes to compare and combine g (≥ 2) independent estimates of kappa. The theory of Section 10.1 applies. Define, for the mth estimate, Vm(κ̂m) to be the squared standard error of κ̂m, that is, the square of the expression in (13.15). The combined estimate of the supposed common value of kappa is, say, κ̂overall = Σ (κ̂m / Vm) / Σ (1 / Vm)… To test the hypothesis that the g underlying values of kappa are equal, the value of χ2equal κ's = Σ (κ̂m − κ̂overall)2 / Vm(κ̂m) may be referred to tables of chi square with g − 1 degrees of freedom. The hypothesis is rejected if the value is significantly large.” — p. 222, equations 13.21 and 13.22 (the sums run over m = 1 to g)
“For testing the hypothesis that the ratings are independent (so that the underlying value of kappa is zero), Fleiss, Cohen, and Everitt (1969) showed that the appropriate standard error of kappa is estimated by…” — p. 219, introducing the null standard error (13.13); the non-null standard error (13.15) follows on p. 220
“Note that the overall value of kappa is equal to the sum of the individual po − pe differences (i.e., of the numerators of the individual kappas) divided by the sum of the individual 1 − pe differences (i.e., of the denominators of the individual kappas)… confirming that κ̂ is a weighted average of the individual κ̂'s.” — p. 221

Verified against the chapter itself, with a check the audit could not do until the book arrived: run on Fleiss's Table 13.1 (three psychiatric diagnoses by two raters, n = 100), the add-in returns kappa 0.6765 and a standard error of 0.0877, against his published .68 and .087; and on the three studies of his Problem 13.3 the tool's common kappa and chi-square agree with equations 13.21 and 13.22 worked by hand to every printed digit. One convention differs, and knowingly: for the test that a single kappa is zero, Fleiss refers the ratio to the simpler null-hypothesis standard error (13.13), one-sided, whereas the tool follows Congalton and Green in using the Fleiss–Cohen–Everitt non-null variance throughout, two-sided. On his example that is 0.076 against 0.087, a Z of 8.9 against 7.7, and the same verdict. Fleiss also passes on the Landis and Koch benchmarks, above .75 excellent and below .40 poor, with the caveat that they hold only “for most purposes.”

Fleiss, Cohen and Everitt (1969)

Fleiss, J. L., J. Cohen, and B. S. Everitt. 1969. Large sample standard errors of kappa and weighted kappa. Psychological Bulletin 72:323–327.

We cite it for: the corrected large-sample variance of kappa and weighted kappa — the θ1 through θ4 formulation shown on the About page, and the variance behind every Z test, confidence interval and comparison printed by the Classification Accuracy (Kappa) and Compare Accuracy Analyses tools.

“Formulas for the standard errors of these two statistics have been given in the literature, but they are in error… Valid formulas for the approximate large-sample variances are given, and their calculation is illustrated using a numerical example.” — Abstract, p. 323

A history worth knowing: even the variance formulas in Cohen's own founding papers were wrong, erring conservative; this correction is the formulation modern software (and this add-in) implements.

Foody (2004)

Foody, G. M. 2004. Thematic map comparison: evaluating the statistical significance of differences in classification accuracy. Photogrammetric Engineering & Remote Sensing 70:627–633. doi.org/10.14358/PERS.70.5.627

We cite it for: the independence caveat on kappa comparisons — the tutorial's honest nuance and the guidance in the Compare Accuracy Analyses tool that same-sample comparisons run conservative and deserve McNemar's paired test.

“The conventional approach to the comparison of kappa coefficients assumes that the samples used in their calculation are independent, an assumption that is commonly unsatisfied because the same sample of ground data sites is often used for each map. Alternative methods to evaluate the statistical significance of differences in accuracy are available for both related and independent samples.” — Abstract, p. 627

Foody (2020)

Foody, G. M. 2020. Explaining the unsuitability of the kappa coefficient in the assessment and comparison of the accuracy of thematic maps obtained by image classification. Remote Sensing of Environment 239:111630. doi.org/10.1016/j.rse.2019.111630

We cite it for: the modern case against kappa as a routine accuracy measure, including the demonstration that a 95%-accurate classification can carry almost any kappa (About — The Kappa debate).

“The kappa coefficient is not an index of accuracy, indeed it is not an index of overall agreement but one of agreement beyond chance. Chance agreement is, however, irrelevant in an accuracy assessment and is anyway inappropriately modelled in the calculation of a kappa coefficient for typical remote sensing applications.” — Abstract
“…for a classification with overall accuracy of 95% the range of possible values of the kappa coefficient is −0.026 to 0.900.” — Abstract
“…researchers are encouraged to provide a set of simple measures and associated outputs such as estimates of per-class accuracy and the confusion matrix when assessing and comparing classification accuracy.” — Abstract

Gwet (2008)

Gwet, K. L. 2008. Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology 61:29–48. doi.org/10.1348/000711006X126600

We cite it for: the AC1 coefficient printed beside kappa in every report of the Classification Accuracy (Kappa) tool, and its robustness to the prevalence paradox (About — the modern statistics).

“It is a fact that these coefficients [pi and kappa] occasionally yield unexpected results in situations known as the paradoxes of kappa. This paper explores the origin of these limitations, and introduces an alternative and more stable agreement coefficient referred to as the AC1 coefficient.” — Abstract, p. 29
“A Monte-Carlo simulation study demonstrates the validity of these variance estimators for confidence interval construction, and confirms the value of AC1 as an improved alternative to existing inter-rater reliability statistics.” — Abstract, p. 29

Kuhn (2008)

Kuhn, M. 2008. Building predictive models in R using the caret package. Journal of Statistical Software 28(5):1–26. doi.org/10.18637/jss.v028.i05

We cite it for: documenting the caret package whose confusionMatrix output the Classification Accuracy (Kappa) tool's per-class suite matches line for line, and which that tool's R snippet uses for verification.

“The caret package, short for classification and regression training, contains numerous tools for developing predictive models using the rich set of models available in R.” — Abstract, p. 1

McNemar (1947)

McNemar, Q. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 12:153–157. doi.org/10.1007/BF02295996

We cite it for: the original paired-proportions test behind the McNemar–Bowker marginal-symmetry statistic, and the correlated-versus-independent-samples distinction behind the independence caveat in the Compare Accuracy Analyses tool.

“Two formulas are presented for judging the significance of the difference between correlated proportions.” — Abstract, p. 153
“There are many situations in which the sampling variance of the difference between two proportions (or percentages) must take into account the fact that the two proportions are not based on independent samples.” — p. 153

Olofsson, Foody, Herold, Stehman, Woodcock and Wulder (2014)

Olofsson, P., G. M. Foody, M. Herold, S. V. Stehman, C. E. Woodcock, and M. A. Wulder. 2014. Good practices for estimating area and assessing accuracy of land change. Remote Sensing of Environment 148:42–57. doi.org/10.1016/j.rse.2014.02.015

We cite it for: the good-practice standard the area-weighted block implements: report the matrix in area proportions, estimate class areas from the reference classification, and attach confidence intervals (About — Class proportions and error-adjusted areas; the Olofsson block of every weighted report from the Classification Accuracy (Kappa) tool).

“(iv) summarize the accuracy assessment by reporting the estimated error matrix in terms of proportion of area and estimates of overall accuracy, user's accuracy (or commission error), and producer's accuracy (or omission error); (v) estimate area of classes… based on the reference classification of the sample units; (vi) quantify uncertainty by reporting confidence intervals for accuracy and area parameters…” — Abstract, p. 42 (recommendations iv–vi of eight)

Pontius and Millones (2011)

Pontius, R. G., Jr., and M. Millones. 2011. Death to Kappa: birth of quantity disagreement and allocation disagreement for accuracy assessment. International Journal of Remote Sensing 32:4407–4429. doi.org/10.1080/01431161.2011.552923

We cite it for: the quantity/allocation partition of total disagreement the report prints, and the sharpest statement of the case against kappa (About — the modern statistics and The Kappa debate).

“This article concludes that these Kappa indices are useless, misleading, and/or flawed for the practical applications in remote sensing that we have seen.” — Abstract

Their recommended replacement — summarizing the cross-tabulation matrix by its quantity disagreement and allocation disagreement components — is exactly the Pontius–Millones line of the report produced by the Classification Accuracy (Kappa) tool.

Sokolova and Lapalme (2009)

Sokolova, M., and G. Lapalme. 2009. A systematic analysis of performance measures for classification tasks. Information Processing and Management 45:427–437. doi.org/10.1016/j.ipm.2009.03.002

We cite it for: formal definitions of the per-class suite (their precision is our user's accuracy, their recall our producer's), the macro-averaging behind Macro-F1, and the invariance finding that F1 never sees correct absences — the reason the Classification Accuracy (Kappa) tool's report pairs it with specificity, negative predictive power and balanced accuracy.

“This paper presents a systematic analysis of twenty four performance measures used in the complete spectrum of Machine Learning classification tasks, i.e., binary, multi-class, multi-labelled, and hierarchical.” — Abstract, p. 427
“Text classification extensively uses Precision and Recall (Sensitivity) which do not detect changes in tn [true negatives]… On contrast, the same invariance for the tn change can be an adversary. Consider the classification of human communication where negative classes are also important.” — p. 431
“Macro-averaging treats all classes equally while micro-averaging favors bigger classes.” — p. 430

A small connection their footnote 3 makes explicit: their single-point “AUC,” ½(sensitivity + specificity), “sometimes referred to as Balanced Accuracy,” is precisely the balanced accuracy of the per-class table.

Stehman and Foody (2019)

Stehman, S. V., and G. M. Foody. 2019. Key issues in rigorous accuracy assessment of land cover products. Remote Sensing of Environment 231:111199. doi.org/10.1016/j.rse.2019.05.018

We cite it for: the current statement of rigorous accuracy-assessment practice, cited beside Olofsson et al. in the Kappa debate as the good-practice standard the Classification Accuracy (Kappa) tool's report follows.

“The article is organized by the three major components of accuracy assessment, the sampling design, response design, and analysis, focusing on good practice methodology that contributes to a rigorous, informative, and honest assessment.” — Abstract
“…documentation of accuracy assessment methods needs to be improved to enhance reproducibility and transparency…” — Abstract — the spirit this Footnotes page tries to honor

Story and Congalton (1986)

Story, M., and R. G. Congalton. 1986. Accuracy assessment: a user's perspective. Photogrammetric Engineering and Remote Sensing 52:397–399.

We cite it for: the producer's-versus-user's accuracy distinction — this three-page brief is where the two accuracies got their names and their interpretations.

“This accuracy value may be referred to as the ‘producer's accuracy,’ because the producer of the classified image/map is interested in how well a specific area on the Earth can be mapped.” — p. 398
“…a better name for this value may be ‘reliability’ (Congalton and Rekas, 1985) or ‘user's accuracy’ because a map user is interested in the reliability of the map, or how well the map represents what is really on the ground.” — p. 398
“…93 percent of the forest has been correctly identified as such, but only 49 percent of those areas identified as forests are actually forests… Although these measures of accuracy may seem very simple, it is critical that they both be considered when assessing the accuracy of a classified image/map.” — pp. 398–399

Their forest example — producer's accuracy 93%, user's accuracy 49%, and a disappointed forester standing in what the map promised was timber — is the direct ancestor of the deciduous-forest walkthrough on the About page.

Thompson (2002)

Thompson, S. K. 2002. Sampling, 2nd ed. John Wiley and Sons, New York.

We cite it for: the finite population correction — the standard sampling-theory shrinkage behind the Estimate Sample Size tool's optional population parameter.

“The quantity (N − n)/N, which may alternatively be written 1 − (n/N), is termed the finite population correction factor.” — p. 15 (§2.2, Estimating the Population Mean)
“In sampling small populations, however, the finite population correction factor may have an appreciable effect in reducing the variance of the estimator, and it is important to include it in the estimate of that variance.” — p. 15

Thompson defines the correction as it acts on a variance; the tool applies its sample-size consequence. Carry the finite-population variance (in its N − 1, hypergeometric form) through the solve-for-n algebra and it yields exactly the tool's n/(1 + (n − 1)/N) — which is also Tortora's equation (2.9); Thompson's (N − n)/N convention yields the near-identical n/(1 + n/N). The two forms differ by less than one sample for any inputs whatever, so the choice is cosmetic; the tool follows Tortora, who in turn credits the correction to Cochran (1963).

Tortora (1978)

Tortora, R. D. 1978. A note on sample size estimation for multinomial populations. The American Statistician 32:100–102. doi.org/10.1080/00031305.1978.10479265

We cite it for: the multinomial sample-size formula behind all three Congalton-Green methods of the Estimate Sample Size tool — the chi-square B, the class-closest-to-50% rule, and the worst-case n = B/4b2.

“A method is described for determining the sample size required for a specified precision simultaneous confidence statement about the parameters of a multinomial population. The method is based on a simultaneous confidence interval procedure due to Goodman, and the results are compared with those obtained by separately considering each cell of the multinomial population as a binomial.” — Abstract, p. 100
“…B is the upper (α/k) × 100th percentile of the χ2 distribution with 1 degree of freedom.” — p. 101 — the Bonferroni share of the alpha that buys the simultaneous guarantee across the classes
“Therefore, one should make k calculations, one for each pair (bi, Πi), i = 1, …, k, and select the largest n as the desired sample size… When bi = b, the only calculation required is for the Πi closest to 1/2. If there is no prior knowledge about the values of the Πi's, a ‘worst’ case calculation of sample size can be made assuming some Πi = 1/2 and bi = b for i = 1, …, k. It is n = B/4b2.” — p. 101 — one paragraph underwriting all three of the tool's multinomial methods: per-class, largest-proportion, and worst-case

His worked example is an anthropologist typing the blood groups A, O, B, and AB on an island: sized one proportion at a time as binomials, the survey needs 377 inhabitants, but the joint statement about all four blood types needs 624 — the price of asking every proportion to land within ±5% simultaneously. And a quiet verification the audit turned up: his finite-population equation (2.9) is algebraically identical to the n/(1 + (n − 1)/N) correction the tool applies.

Van Rijsbergen (1979)

Van Rijsbergen, C. J. 1979. Information Retrieval, 2nd ed. Butterworths, London. Freely available from the author at dcs.gla.ac.uk/Keith.

We cite it for: the origin of the F-measure: Chapter 7 (Evaluation) develops the effectiveness measure E from precision and recall, and F = 1 − E at equal weighting is the F1 the Classification Accuracy (Kappa) tool's report prints per class.

“It is recall and precision which attempt to measure what is now known as the effectiveness of the retrieval system. In other words it is a measure of the ability of the system to retrieve relevant documents while at the same time holding back non-relevant one[s].” — Chapter 7

Wynne (2003)

Wynne, J. J. 2003. Landscape-scale modeling of vegetation land cover and songbird habitat, Pinaleños Mountains, Arizona. M.S. thesis, Northern Arizona University, Flagstaff.

We cite it for: the Mexican Jay case study — the competing classification-tree and literature-based habitat models, assessed with Cohen's Kappa and chosen by multi-criteria selection when significance was out of reach.

“Using a competing models framework, I modeled habitat at the landscape-scale using classification tree and logistic regression models for eight songbird species on the Pinaleños Mountains, southeastern Arizona. Classification tree output and literature-derived information were used for creating predicted distribution maps with a GIS, accuracy assessed using the 2002 dataset and a Cohen's Kappa, and selected using a multi-criteria selection approach.” — Chapter 3 abstract, p. 52

Deliberately exempt

Every cited work above has now been checked against the original text. Two works are deliberately exempt: Congalton and Green (1999), whose influence on the original tool was so pervasive that nearly the whole book applies rather than any quotable passage, and Jenness and Wynne (2007), which is not a source for this tool but its direct ancestor.