Mahalanobis Distance Raster

Home Range · geoprocessing tool · by Jeff Jenness
Works at every ArcGIS Pro license level
Learn more About Mahalanobis distances covers the theory — what the distances measure, a worked example, the chi-square p-value transform, and how the four Mahalanobis tools fit together.

Summary

Creates a raster of Mahalanobis distances measuring, for every cell, how far that cell's combination of landscape-variable values lies from a multivariate reference mean, scaled by the inverse covariance among the variables. Low values mark ground most similar to the reference conditions. If those reference conditions represent, say, the ideal habitat for some wildlife species, then the Mahalanobis raster tells you how similar every point on the ground is to that ideal habitat — a great first step toward identifying critical habitat or steering management actions. This is a similarity map of the whole landscape, and the classic multivariate habitat-modeling surface of Clark et al. (1993) and Knick and Dyer (1997).

How it works

The Home Range Tools gallery open on the ribbon, with the Mahalanobis Distance Raster button, in the Mahalanobis Tools row, outlined in blue
Where to find it: Mahalanobis Distance Raster is in the Mahalanobis Tools row of the Home Range Tools gallery, in the Home Range group of the Wildlife and Forestry tab.

The squared Mahalanobis distance for a cell is the quadratic form D² = (x − μ)′ Σ⁻¹ (x − μ), where x is the cell's vector of landscape-variable values, μ is the mean vector, and Σ⁻¹ is the inverse covariance matrix (bold indicates a vector or matrix, non-bold a single number). The computation processes the rasters in manageable blocks of cells rather than loading everything into memory at once, so it won't run out of memory no matter how large the rasters are or how many variables you use — and it needs no Spatial Analyst license.

The landscape variables

Variables can be rasters (cell values are always used, never attribute-table values) and/or polygon feature classes with a numeric field, which are rasterized to the analysis grid. A multiband raster automatically expands to one variable per band, and a Bands to exclude table lets you drop specific bands. At least one raster is required — it sets the analysis grid's cell size, coordinate system, and extent (adjustable through the environments).

Order matters The variables become the columns of the mean vector and covariance matrix in the order listed (rasters first, then polygons). When you reuse a saved mean and covariance, the same datasets must be supplied in the same order — a different order silently produces incorrect distances. We don't want our elevation values evaluated against our slope mean and variance! The tools in this add-in help you avoid this problem by writing the variable order into the saved tables' metadata, so the other tools know automatically what order to analyze the variables in. If you use statistical tables generated from other sources, you'll need to take care to put the variables in the correct order yourself.

Three sources for the reference statistics

Calculate directly from sample point features — the common wildlife case: the landscape variables are sampled at your animal locations (optionally only the selected ones), and the mean vector and covariance matrix are computed from those samples. Points outside the landscape data, or on NoData cells, are excluded and reported.

Categorical raster — the reference sample is every cell of a categorical raster belonging to a chosen category. If the raster has an attribute table, pick any of its fields as the Category attribute field and choose the category from that field's values — the tool then gathers every matching cell value automatically. Many classified rasters describe the same cells at several resolutions: a Landfire EVT raster, for example, carries both the fine-grained EVT_NAME (“Rocky Mountain Aspen Forest and Woodland”) and the broad EVT_LF (“Tree”), and choosing EVT_LF = Tree sweeps in every one of the dozens of tree-type cell values in one step. Leave the field blank to enter the numeric cell value directly instead — the right choice for a raster with no attribute table. Scripts and models may use either form: pass the field and category value exactly as the dialog does, or pass literal cell values (see the Python section).

Existing tables — a previously saved mean vector and covariance matrix, from this tool's optional table outputs or from Calculate Statistical Matrices. The variable order is validated against the tables' stored order.

The Cell value sampling option controls how raster values are read at sample points (and how rasters that don't align to the analysis grid are resampled): Exact (nearest cell) or Interpolated (bilinear). Bilinear is more appropriate for continuous surfaces like elevation; exact is right for categorical rasters — interpolation should never be used on nominal (unordered) categories.

Three flavors of output

All three are functions of the same per-cell squared distance:

Output valueWhat it is
Squared distance (D²) The quadratic form itself — the quantity traditionally mapped in ecological and habitat Mahalanobis models, and the input required by the chi-square transform.
Distance (D) The square root of D²: the literal Mahalanobis distance, analogous to Euclidean distance.
Chi-square p-value The upper-tail probability P(χ² > D²) with degrees of freedom equal to the number of variables: a 0 to 1 similarity score where values near 1 indicate a cell closely matching the reference conditions. See the degrees-of-freedom note on the About page.

Inverting the covariance

Standard (inverse) is the ordinary matrix inverse — fast, exact, and the right choice in almost all cases. It fails if the covariance matrix is singular, which usually means a variable has no variance (all values identical) or two variables are perfectly correlated. Pseudo-inverse (SVD), the Moore-Penrose pseudo-inverse, returns a stable least-squares inverse instead of failing when variables are strongly collinear — though it can mask redundant variables. The tool reports the covariance condition number; very large values indicate nearly collinear variables (the correlation-matrix output is the natural way to hunt down which ones).

Saving the statistics for reuse

Four optional checkboxes also write the computed mean vector, covariance, inverse covariance, and correlation tables, auto-named from the output raster. The mean and covariance pair is what the Existing tables source reads back — compute the statistics once from your best sample, then apply them to other landscapes, other years, or the at Points tool, without resampling.

Environments

The tool honors the Output Coordinate System, Cell Size, Processing Extent, Snap Raster, and Mask environments, which together control the analysis grid. Other geoprocessing environments do not affect it.

Scenario 1: statistics from animal locations

The classic wildlife case. We have locations recorded for some animal we are interested in, and we can estimate the “ideal” habitat as that which is most similar to the mean elevation, slope and landscape curvature observed at those animal locations. Just as important, the model also takes advantage of the covariance among these variables at those locations: if our animal tends to use steep slopes at low elevations and shallow slopes higher up, that relationship is baked into the covariance matrix, and the distances honor it — a cell with a typical elevation and a typical slope, but in an atypical combination, still scores as dissimilar.

Three stacked map panels of the same landscape, curvature, slope in degrees, and elevation in meters, each overlaid with the same blue animal location points
The inputs: elevation, slope and curvature rasters, with the animal locations overlaid.
Important assumptions Two things must be true for this surface to mean what we want it to mean: that these points represent a good sample of the animal's use of the landscape (enough of them, collected without strong bias toward roads, daytime, or easy terrain), and that these variables are actually important to the animal. The tool will happily compute distances from any variables you give it; only your professional expertise can say whether elevation, slope and curvature are what this species responds to.
The Mahalanobis Distance Raster pane with elev_clip, slope_clip and curvature as raster variables, statistics calculated directly from the Sample_Points_for_Mahal layer, bilinear sampling, and a squared-distance output
The dialog: three raster variables, statistics calculated directly from the sample point features, D² out.
The output Mahalanobis D-squared surface, blue similar terrain enveloping the animal locations, red dissimilar terrain along the canyon bottoms and washes
The output D² surface with the sample points overlaid: blue (low D²) marks terrain most similar to the conditions at the animal locations; the canyon bottoms and washes score as most dissimilar.

Points that don't lie within the analysis area are handled for you: any point falling outside the extent of one or more landscape rasters (or on a NoData cell) is automatically excluded from the statistics, and the analysis report tells you exactly how many were dropped and how many were used:

The geoprocessing report with a highlighted warning reading 2 of 48 sample points fell outside the extent of one or more landscape rasters and were excluded from the statistics, followed by 46 of 48 sample points were used
Off-extent points are excluded automatically and reported — here, 2 of 48 fell outside the landscape rasters and the statistics used the remaining 46.

Scenario 2: statistics from a categorical raster

Now suppose that instead of animal locations we have a categorical raster that includes a category we care about. Here we use the Landfire Existing Vegetation Type dataset, symbolized by its EVT subclass field, and ask how similar the landscape is to the ground classified as Evergreen open tree canopy — calculating D² from the same three variables as before. The reference sample is every cell of that category (the analysis report tells you how many cells of the categorical raster matched), and the mean and covariance describe the terrain that vegetation type occupies.

The Landfire EVT subclass map of the study area, a mosaic of vegetation categories including several evergreen tree canopy classes, shrubland, and grassland types
The Landfire EVT raster, symbolized by subclass — Evergreen open tree canopy is one of its categories.
The Mahalanobis D-squared surface built from the Evergreen open tree canopy cells, most of the uplands scoring similar in blue, washes and rocky breaks scoring dissimilar
D² from the Evergreen-open-tree-canopy cells: the surface maps how similar every cell's terrain is to the terrain that category occupies.
Important assumptions

The first assumption in this example is that Landfire is categorized correctly — and you should weigh this same assumption yourself if you are using a similarly-modeled categorical raster of your own. It is a bigger assumption than it sounds, because Landfire's vegetation map is itself a model: a remote-sensing classification with its own error rate, which varies from class to class. That makes our Mahalanobis surface a second-level model, a model built on top of another model, and errors propagate quickly when you stack models this way. Every misclassified cell smuggles the wrong terrain into our reference sample, nudging the mean and covariance away from the true conditions of the vegetation type — and then every cell of the output inherits statistics contaminated by those errors. The reliabilities compound roughly multiplicatively: a classification that is 80% right, feeding a habitat model that would be 80% faithful given perfect input, leaves something closer to a 64% model, not an 80% one. None of this makes the approach invalid — it makes the output an approximation whose uncertainty is inherited from Landfire's, and worth interpreting with that in mind.

The second assumption is that the landscape variables we chose are relevant to this vegetation type, and sufficient to model it. Elevation, slope and curvature say nothing about soils, precipitation, aspect, or fire history — all of which shape where evergreen canopy actually grows. Should we have included more or different variables? That question has no automatic answer; it is an ecological judgment the tool cannot make for you.

Spatial autocorrelation, and what it means for interpretation A categorical reference sample can easily contain hundreds of thousands of cells, but they are not hundreds of thousands of independent samples. Neighboring raster cells are strongly spatially autocorrelated — elevation and slope vary smoothly, and vegetation maps come in patches, so each cell largely repeats the information of its neighbors. A cluster of raster cells does not represent independent samples, and the effective sample size is far smaller than the cell count. For the descriptive use these tools make of the sample — the mean and covariance of the terrain that category occupies within the analysis area — this is no great harm: we are summarizing conditions, close to a census of the category, not inferring from a random draw. The caution comes in interpreting the output. The chi-square p-values should be read as a similarity score — a convenient 0-to-1 rescaling — and never as statistical significance backed by an enormous sample, because that sample is nowhere near as informative as its cell count suggests. The same warning applies downstream: any analysis that treats individual cells of the output as independent observations (accuracy tests, cell-counting comparisons between models) will overstate its confidence. Read the surface as a ranking of the landscape from most to least similar — that ranking is exactly what the Mahalanobis distance delivers, autocorrelation or no.

ModelBuilder

The output raster chains onward like any raster — to the Chi-Square Transform for a 0 to 1 similarity-score surface, to a Reclassify for candidate-habitat classes, or into corridor modeling. The optional statistics tables are also model outputs, so one model can compute the statistics and fan them out to several downstream Mahalanobis runs.

A ModelBuilder diagram: Elevation, Slope, Curvature and Landfire rasters feeding the Mahalanobis Distance Raster tool, which outputs the Mahalanobis_Landfire raster plus the Means, Covariance, Inverse Covariance and Correlation tables
The tool in a model — the categorical scenario above as a diagram: elevation, slope, curvature and the Landfire raster in; the Mahalanobis raster and all four statistics tables out, each ready to chain onward.

Parameters

LabelExplanationData type
Raster Landscape VariablesRequired · raster_vars The raster datasets defining the multivariate space; cell values are used, never attribute-table values. Multiband rasters expand to one variable per band. At least one raster is required. Raster Layer (multiple)
Polygon Landscape VariablesOptional · poly_vars Rows of polygon feature class + numeric field, rasterized to the analysis grid. Polygon variables follow the rasters in covariance order. Value Table
Bands to excludeOptional · exclude_bands Rows of multiband raster + 1-based band number to drop. Value Table
Source of mean vector and covariance matrixRequired · stats_source Sample point features, categorical raster, or existing tables. Choose the source first; the relevant inputs below enable accordingly. String
Sample pointsOptional · points The reference points (point-features source only). Feature Layer
Use only the selected pointsOptional · use_selected Restrict the statistics to the layer's active selection. Boolean
Categorical rasterOptional · cat_raster The raster whose matching cells form the reference sample (categorical source only). Raster Layer
Category attribute fieldOptional · cat_field A field from the raster's attribute table naming the categories the way you want to choose them (Landfire's EVT_NAME or EVT_LF, say); leave blank to enter a literal cell value instead. Field
Category valueOptional · cat_value The category whose cells form the reference sample — a dropdown of the chosen field's values, with every matching cell value gathered automatically. With no field, literal cell value(s), semicolon-separated for several (“3;4;5”). String
Mean vector tableOptional · mean_table A saved mean vector (existing-tables source only). Table View
Covariance matrix tableOptional · cov_table A saved covariance matrix; its stored variable order validates the variable list. Table View
Cell value samplingOptional · sample_method Exact (nearest cell) or Interpolated (bilinear); see above. String
Covariance inversion methodRequired · inversion Standard (inverse) or Pseudo-inverse (SVD); see above. String
Output Mahalanobis rasterRequired · out_raster The output raster; defaults to Mahalanobis in the current workspace, auto-incremented. Raster Dataset
Output valueRequired · out_type Squared distance (D²), Distance (D), or Chi-square p-value. String
Generate / save statistics tablesOptional · gen_mean…save_corr Checkbox-and-output pairs for the mean vector, covariance, inverse covariance, and correlation tables, auto-named from the raster name (_Means, _Covariance, _Inv_Covariance, _Correlation). Boolean + Table

Python

A D² raster from three rasters in the active map, with statistics from a sample-point layer, saving all four statistics tables:

import arcpy
arcpy.ImportToolbox(r"C:\path\to\JennessEnterprisesTools.pyt")  # your install path
out_gdb = r"C:\Project\Mahalanobis.gdb"
arcpy.jenness.MahalanobisRaster(
    raster_vars="elev_clip;slope_clip;curvature",
    sample_method="Interpolated (bilinear)",
    stats_source="Calculate directly from sample point features",
    points="Sample_Points_for_Mahal",
    inversion="Standard (inverse)",
    out_raster=out_gdb + r"\Mahalanobis",
    out_type="Squared distance (D2): Traditional for ecological modeling",
    gen_mean=True,  save_mean=out_gdb + r"\Mahalanobis_Means",
    gen_cov=True,   save_cov=out_gdb + r"\Mahalanobis_Covariance",
    gen_invcov=True, save_invcov=out_gdb + r"\Mahalanobis_Inv_Covariance",
    gen_corr=True,  save_corr=out_gdb + r"\Mahalanobis_Correlation")

Categorical-raster statistics work the same way in a script as in the dialog: pass the attribute field and category value, and the tool gathers every matching cell value itself — the script never needs to look up the codes:

arcpy.jenness.MahalanobisRaster(
    raster_vars="elev_clip;slope_clip;curvature",
    stats_source="Categorical raster",
    cat_raster="Landfire",
    cat_field="EVT_LF",
    cat_value="Tree",
    inversion="Standard (inverse)",
    out_raster=out_gdb + r"\Mahalanobis_Tree",
    out_type="Squared distance (D2): Traditional for ecological modeling")

Alternatively, skip the field and supply literal cell value(s) yourself — a single value, or several separated by semicolons, every listed value joining the reference sample:

    # ...same call, but with:
    cat_field=None,
    cat_value="7016;7019;7023",

Recommended citation

Jenness, J. 2026. Mahalanobis Distance Raster. Wildlife and Forestry Tools add-in for ArcGIS Pro, v. 1.99 (September 2026). Jenness Enterprises. Available at: https://github.com/JeffJenness/Wildlife_Tools.

Credits and references

By Jeff Jenness, Jenness Enterprises (www.jennessent.com), ported from his ArcView Mahalanobis Distances extension and the ArcMap Land Facet Corridor Tools.

Licensing information

Works at every ArcGIS Pro license level (Basic, Standard, Advanced). No extension licenses are required — including Spatial Analyst.