Rev 52927 | Blame | Compare with Previous | Last modification | View Log | Download | RSS feed
% File src/library/base/man/agrep.Rd% Part of the R package, http://www.R-project.org% Copyright 1995-2011 R Core Development Team% Distributed under GPL 2 or later\name{agrep}\alias{agrep}\alias{fuzzy matching}\title{Approximate String Matching (Fuzzy Matching)}\description{Searches for approximate matches to \code{pattern} (the first argument)within each element of the string \code{x} (the second argument) usingthe Levenshtein edit distance.}\usage{agrep(pattern, x, ignore.case = FALSE, value = FALSE,max.distance = 0.1, useBytes = FALSE)}\arguments{\item{pattern}{a non-empty character string to be matched (\emph{not}a regular expression!). Coerced by \code{as.character} to a stringif possible.}\item{x}{character vector where matches are sought. Coerced by\code{as.character} to a character vector if possible.}\item{ignore.case}{if \code{FALSE}, the pattern matching is \emph{casesensitive} and if \code{TRUE}, case is ignored during matching.}\item{value}{if \code{FALSE}, a vector containing the (integer)indices of the matches determined is returned and if \code{TRUE}, avector containing the matching elements themselves is returned.}\item{max.distance}{Maximum distance allowed for a match. Expressedeither as integer, or as a fraction of the \emph{pattern} length (will bereplaced by the smallest integer not less than the correspondingfraction of the pattern length), or a list with possible components\describe{\item{\code{all}:}{maximal (overall) distance}\item{\code{insertions}:}{maximum number/fraction of insertions}\item{\code{deletions}:}{maximum number/fraction of deletions}\item{\code{substitutions}:}{maximum number/fraction ofsubstitutions}}If \code{all} is missing, it is set to 10\%, the other componentsdefault to \code{all}. The component names can be abbreviated.}\item{useBytes}{logical. in a multibyte locale, should the comparisonbe character-by-character (the default) or byte-by-byte.}}\details{The Levenshtein edit distance is used as measure of approximateness:it is the total number of insertions, deletions and substitutionsrequired to transform one string into another.As from \R 2.10.0 this uses \code{tre} by Ville Laurikari(\url{http://http://laurikari.net/tre/}), which supports MBCScharacter matching much better than the previous version.}\note{Since someone who read the description carelessly even filed a bugreport on it, do note that this matches substrings of each element of\code{x} (just as \code{\link{grep}} does) and \bold{not} wholeelements.}\value{Either a vector giving the indices of the elements that yielded amatch, or, if \code{value} is \code{TRUE}, the matched elements (aftercoercion, preserving names but no other attributes).}\author{Original version by David Meyer. Current version by Brian Ripley.}\seealso{\code{\link{grep}}}\examples{agrep("lasy", "1 lazy 2")agrep("lasy", c(" 1 lazy 2", "1 lasy 2"), max = list(sub = 0))agrep("laysy", c("1 lazy", "1", "1 LAZY"), max = 2)agrep("laysy", c("1 lazy", "1", "1 LAZY"), max = 2, value = TRUE)agrep("laysy", c("1 lazy", "1", "1 LAZY"), max = 2, ignore.case = TRUE)}\keyword{character}