Rev 82028 | Blame | Compare with Previous | Last modification | View Log | Download | RSS feed
% File src/library/stats/man/ks.test.Rd% Part of the R package, https://www.R-project.org% Copyright 1995-2022 R Core Team% Distributed under GPL 2 or later\name{ks.test}\alias{ks.test}\alias{ks.test.default}\alias{ks.test.formula}\encoding{UTF-8}\title{Kolmogorov-Smirnov Tests}\description{Perform a one- or two-sample Kolmogorov-Smirnov test.}\usage{ks.test(x, \dots)\method{ks.test}{default}(x, y, \dots,alternative = c("two.sided", "less", "greater"),exact = NULL, simulate.p.value = FALSE, B = 2000)\method{ks.test}{formula}(formula, data, subset, na.action, \dots)}\arguments{\item{x}{a numeric vector of data values.}\item{y}{either a numeric vector of data values, or a character stringnaming a cumulative distribution function or an actual cumulativedistribution function such as \code{pnorm}. Only continuous CDFsare valid.}\item{\dots}{for the default method, parameters of the distributionspecified (as a character string) by \code{y}. Otherwise, furtherarguments to be passed to or from methods.}\item{alternative}{indicates the alternative hypothesis and must beone of \code{"two.sided"} (default), \code{"less"}, or\code{"greater"}. You can specify just the initial letter of thevalue, but the argument name must be given in full.See \sQuote{Details} for the meanings of the possible values.}\item{exact}{\code{NULL} or a logical indicating whether an exactp-value should be computed. See \sQuote{Details} for the meaning of\code{NULL}.}\item{simulate.p.value}{a logical indicating whether to computep-values by Monte Carlo simulation.}\item{B}{an integer specifying the number of replicates used in theMonte Carlo test.}\item{formula}{a formula of the form \code{lhs ~ rhs} where \code{lhs}is a numeric variable giving the data values and \code{rhs} either\code{1} for a one-sample test or a factor with two levels givingthe corresponding groups for a two-sample test.}\item{data}{an optional matrix or data frame (or similar: see\code{\link{model.frame}}) containing the variables in theformula \code{formula}. By default the variables are taken from\code{environment(formula)}.}\item{subset}{an optional vector specifying a subset of observationsto be used.}\item{na.action}{a function which indicates what should happen whenthe data contain \code{NA}s. Defaults to\code{getOption("na.action")}.}}\details{If \code{y} is numeric, a two-sample (Smirnov) test of the null hypothesisthat \code{x} and \code{y} were drawn from the same \emph{continuous}distribution is performed.Alternatively, \code{y} can be a character string naming a continuous(cumulative) distribution function, or such a function. In this case,a one-sample (Kolmogorov) test is carried out of the null that the distributionfunction which generated \code{x} is distribution \code{y} withparameters specified by \code{\dots}.The presence of ties always generates a warning, since continuousdistributions do not generate them. If the ties arose from roundingthe tests may be approximately valid, but even modest amounts ofrounding can have a significant effect on the calculated statistic.Missing values are silently omitted from \code{x} and (in thetwo-sample case) \code{y}.The possible values \code{"two.sided"}, \code{"less"} and\code{"greater"} of \code{alternative} specify the null hypothesisthat the true distribution function of \code{x} is equal to, not lessthan or not greater than the hypothesized distribution function(one-sample case) or the distribution function of \code{y} (two-samplecase), respectively. This is a comparison of cumulative distributionfunctions, and the test statistic is the maximum difference in value,with the statistic in the \code{"greater"} alternative being\eqn{D^+ = \max_u [ F_x(u) - F_y(u) ]}{D^+ = max[F_x(u) - F_y(u)]}.Thus in the two-sample case \code{alternative = "greater"} includesdistributions for which \code{x} is stochastically \emph{smaller} than\code{y} (the CDF of \code{x} lies above and hence to the left of thatfor \code{y}), in contrast to \code{\link{t.test}} or\code{\link{wilcox.test}}.Exact p-values are not available for the one-sample case in thepresence of ties.If \code{exact = NULL} (the default), anexact p-value is computed if the sample size is less than 100 in theone-sample case \emph{and there are no ties}, and if the product ofthe sample sizes is less than 10000 in the two-sample case, with orwithout ties (using the algorithm described in Schröer and Trenkler, 1995).Otherwise, asymptotic distributions are used whose approximations maybe inaccurate in small samples. In the one-sample two-sided case,exact p-values are obtained as described in Marsaglia, Tsang & Wang(2003) (but not using the optional approximation in the right tail, sothis can be slow for small p-values). The formula of Birnbaum &Tingey (1951) is used for the one-sample one-sided case.If a one-sample test is used, the parameters specified in\code{\dots} must be pre-specified and not estimated from the data.There is some more refined distribution theory for the KS test withestimated parameters (see Durbin, 1973), but that is not implementedin \code{ks.test}.}\value{A list inheriting from classes \code{"ks.test"} and \code{"htest"}containing the following components:\item{statistic}{the value of the test statistic.}\item{p.value}{the p-value of the test.}\item{alternative}{a character string describing the alternativehypothesis.}\item{method}{a character string indicating what type of test wasperformed.}\item{data.name}{a character string giving the name(s) of the data.}}\source{The two-sided one-sample distribution comes \emph{via}Marsaglia, Tsang and Wang (2003).Exact distributions for the two-sample (Smirnov) test are computedby the algorithm proposed Schröer (1991) and Schröer & Trenkler (1995).}\references{Z. W. Birnbaum and Fred H. Tingey (1951).One-sided confidence contours for probability distribution functions.\emph{The Annals of Mathematical Statistics}, \bold{22}/4, 592--596.\doi{10.1214/aoms/1177729550}.William J. Conover (1971).\emph{Practical Nonparametric Statistics}.New York: John Wiley & Sons.Pages 295--301 (one-sample Kolmogorov test),309--314 (two-sample Smirnov test).Durbin, J. (1973).\emph{Distribution theory for tests based on the sample distributionfunction}.SIAM.W. Feller (1948).On the Kolmogorov-Smirnov limit theorems for empirical distributions.\emph{The Annals of Mathematical Statistics}, \bold{19}(2), 177--189.\doi{10.1214/aoms/1177730243}.George Marsaglia, Wai Wan Tsang and Jingbo Wang (2003).Evaluating Kolmogorov's distribution.\emph{Journal of Statistical Software}, \bold{8}/18.\doi{10.18637/jss.v008.i18}.Gunar Schröer (1991),Computergestützte statistische Inferenz am Beispiel derKolmogorov-Smirnov Tests.Diplomarbeit Universität Osnabrück.Gunar Schröer and Dietrich Trenkler (1995).Exact and Randomization Distributions of Kolmogorov-Smirnov Tests forTwo or Three Samples.\emph{Computational Statistics & Data Analysis}, \bold{20}(2),185--202.\doi{10.1016/0167-9473(94)00040-P}.}\seealso{\code{\link{psmirnov}}.\code{\link{shapiro.test}} which performs the Shapiro-Wilk test fornormality.}\examples{require("graphics")x <- rnorm(50)y <- runif(30)# Do x and y come from the same distribution?ks.test(x, y)# Does x come from a shifted gamma distribution with shape 3 and rate 2?ks.test(x+2, "pgamma", 3, 2) # two-sided, exactks.test(x+2, "pgamma", 3, 2, exact = FALSE)ks.test(x+2, "pgamma", 3, 2, alternative = "gr")# test if x is stochastically larger than x2x2 <- rnorm(50, -1)plot(ecdf(x), xlim = range(c(x, x2)))plot(ecdf(x2), add = TRUE, lty = "dashed")t.test(x, x2, alternative = "g")wilcox.test(x, x2, alternative = "g")ks.test(x, x2, alternative = "l")# with ties, example from Schröer and Trenkler (1995)# D = 3 / 7, p = 0.2424242ks.test(c(1, 2, 2, 3, 3), c(1, 2, 3, 3, 4, 5, 6), exact = TRUE)# formula interface, see ?wilcox.testks.test(Ozone ~ Month, data = airquality,subset = Month \%in\% c(5, 8))}\keyword{htest}