Rev 71259 | Blame | Compare with Previous | Last modification | View Log | Download | RSS feed
% File src/library/base/man/connections.Rd% Part of the R package, https://www.R-project.org% Copyright 1995-2016 R Core Team% Distributed under GPL 2 or later\name{connections}\title{Functions to Manipulate Connections (Files, URLs, ...)}\alias{connections}\alias{connection}\alias{file}\alias{clipboard}\alias{pipe}\alias{fifo}\alias{gzfile}\alias{unz}\alias{bzfile}\alias{xzfile}\alias{url}\alias{socketConnection}\alias{open}\alias{open.connection}\alias{isOpen}\alias{isIncomplete}\alias{close}\alias{close.connection}\alias{flush}\alias{flush.connection}\alias{print.connection}\alias{summary.connection}\concept{encoding}\concept{compression}\concept{decompression}\concept{zip}\concept{gzip}\concept{bzip2}\concept{lzma}\description{Functions to create, open and close connections, i.e.,\dQuote{generalized files}, such as possibly compressed files, URLs,pipes, etc.}\usage{file(description = "", open = "", blocking = TRUE,encoding = getOption("encoding"), raw = FALSE,method = getOption("url.method", "default"))url(description, open = "", blocking = TRUE,encoding = getOption("encoding"),method = getOption("url.method", "default"))gzfile(description, open = "", encoding = getOption("encoding"),compression = 6)bzfile(description, open = "", encoding = getOption("encoding"),compression = 9)xzfile(description, open = "", encoding = getOption("encoding"),compression = 6)unz(description, filename, open = "", encoding = getOption("encoding"))pipe(description, open = "", encoding = getOption("encoding"))fifo(description, open = "", blocking = FALSE,encoding = getOption("encoding"))socketConnection(host = "localhost", port, server = FALSE,blocking = FALSE, open = "a+",encoding = getOption("encoding"),timeout = getOption("timeout"))open(con, \dots)\method{open}{connection}(con, open = "r", blocking = TRUE, \dots)close(con, \dots)\method{close}{connection}(con, type = "rw", \dots)flush(con)isOpen(con, rw = "")isIncomplete(con)}\arguments{\item{description}{character string. A description of the connection:see \sQuote{Details}.}\item{open}{character string. A description of how to open the connection(if it should be opened initially). See section \sQuote{Modes} forpossible values.}\item{blocking}{logical. See the \sQuote{Blocking} section.}\item{encoding}{The name of the encoding to be assumed. See the\sQuote{Encoding} section.}\item{raw}{logical. If true, a \sQuote{raw} interface is used whichwill be more suitable for arguments which are not regular files,e.g.\sspace{}character devices. This suppresses the check for a compressedfile when opening for text-mode reading, and asserts that the\sQuote{file} may not be seekable.}\item{method}{character string, partially matched to\code{c("default", "internal", "wininet", "libcurl")}:%% FIXME: Consider "auto", as in download.file()see \sQuote{Details}.}\item{compression}{integer in 0--9. The amount of compression to beapplied when writing, from none to maximal available. For\code{xzfile} can also be negative: see the \sQuote{Compression}section.}\item{timeout}{numeric: the timeout (in seconds) to be used for thisconnection. Beware that some OSes may treat very large values aszero: however the POSIX standard requires values up to 31 days to besupported.}\item{filename}{a filename within a zip file.}\item{host}{character string. Host name for the port.}\item{port}{integer. The TCP port number.}\item{server}{logical. Should the socket be a client or a server?}\item{con}{a connection.}\item{type}{character string. Currently ignored.}\item{rw}{character string. Empty or \code{"read"} or \code{"write"},partial matches allowed.}\item{\dots}{arguments passed to or from other methods.}}\details{The first nine functions create connections. By default theconnection is not opened (except for a \code{socketConnection}), but maybe opened by setting a non-empty value of argument \code{open}.For \code{file} the description is a path to the file to be opened ora complete URL (when it is the same as calling \code{url}), or\code{""} (the default) or \code{"clipboard"} (see the\sQuote{Clipboard} section). Use \code{"stdin"} to refer to theC-level \sQuote{standard input} of the process (which need not beconnected to anything in a console or embedded version of \R, and isnot in \code{RGui} on Windows). See also \code{\link{stdin}()} forthe subtly different R-level concept of \code{stdin}.For \code{url} the description is a complete URL including scheme(such as \samp{http://}, \samp{https://}, \samp{ftp://} or\samp{file://}). Method \code{"internal"} is that available sinceconnections were introduced, method \code{"wininet"} is only availableon Windows (it uses the WinINet functions of that OS) and method\code{"libcurl"} (using the library of that name:\url{http://curl.haxx.se/libcurl/}) is required on a Unix-alike butoptional on Windows. Method \code{"default"} uses method\code{"internal"} for \samp{file:} URLs and \code{"libcurl"} for\code{ftps:} URLs. On a Unix-alike it uses \code{"internal"} for\samp{http:} and \code{ftp:} URLs and \code{"libcurl"} for\samp{https:} URLs; on Windows \code{"wininet"} for \samp{http:},\code{ftp:} and \samp{https:} URLs. Proxies can be specified: see\code{\link{download.file}}.For \code{gzfile} the description is the path to a file compressed by\command{gzip}: it can also open for reading uncompressed files andthose compressed by \command{bzip2}, \command{xz} or \command{lzma}.For \code{bzfile} the description is the path to a file compressed by\command{bzip2}.For \code{xzfile} the description is the path to a file compressed by\command{xz} (\url{https://en.wikipedia.org/wiki/Xz}) or (for readingonly) \command{lzma} (\url{https://en.wikipedia.org/wiki/LZMA}).\code{unz} reads (only) single files within zip files, in binary mode.The description is the full path to the zip file, with \file{.zip}extension if required.For \code{pipe} the description is the command line to be piped to orfrom. This is run in a shell, on Windows that specified by the\env{COMSPEC} environment variable.For \code{fifo} the description is the path of the fifo. (Support for\code{fifo} connections is optional but they are available on mostUnix platforms and on Windows.)The intention is that \code{file} and \code{gzfile} can be usedgenerally for text input (from files, \samp{http://} and\samp{https://} URLs) and binary input respectively.\code{open}, \code{close} and \code{seek} are generic functions: thefollowing applies to the methods relevant to connections.\code{open} opens a connection. In general functions usingconnections will open them if they are not open, but then close themagain, so to leave a connection open call \code{open} explicitly.\code{close} closes and destroys a connection. This will happenautomatically in due course (with a warning) if there is no longer an\R object referring to the connection.A maximum of 128 connections can be allocated (not necessarily open)at any one time. Three of these are pre-allocated (see\code{\link{stdout}}). The OS will impose limits on the numbers ofconnections of various types, but these are usually larger than 125.\code{flush} flushes the output stream of a connection open forwrite/append (where implemented, currently for file and clipboardconnections, \code{\link{stdout}} and \code{\link{stderr}}).If for a \code{file} or (on most platforms) a \code{fifo} connectionthe description is \code{""}, the file/fifo is immediately opened (in\code{"w+"} mode unless \code{open = "w+b"} is specified) and unlinkedfrom the file system. This provides a temporary file/fifo to write toand then read from.}\section{URLs}{\code{url} and \code{file} support URL schemes \samp{file://},\samp{http://}, \samp{https://} and \samp{ftp://}.\code{method = "libcurl"} allows more schemes: exactly which schemesis platform-dependent (see \code{\link{libcurlVersion}}), but allUnix-alike platforms will support \samp{https://} and most platformswill support \samp{ftps://}.Most methods do not percent-encode special characters such as spacesin \samp{http://} URLs (see \code{\link{URLencode}}), but it seems the\code{"wininet"} method does.A note on \samp{file://} URLs. The most general form (from RFC1738) is\samp{file://host/path/to/file}, but \R only accepts the form with anempty \code{host} field referring to the local machine.On a Unix-alike, this is then \samp{file:///path/to/file}, where\samp{path/to/file} is relative to \file{/}. So although the thirdslash is strictly part of the specification not part of the path, thiscan be regarded as a way to specify the file \file{/path/to/file}. Itis not possible to specify a relative path using a file URL.In this form the path is relative to the root of the filesystem, not aWindows concept. The standard form on Windows is\samp{file:///d:/R/repos}: for compatibility with earlier versions of\R and Unix versions, any other form is parsed as \R as \samp{file://}plus \code{path_to_file}. Also, backslashes are accepted within thepath even though RFC1738 does not allow them.No attempt is made to decode a percent-encoded \samp{file:} URL: call\code{\link{URLdecode}} if necessary.The \code{"internal"} method does not follow re-directed HTTP URLs:both methods \code{"wininet"} (the default on Windows) and\code{"libcurl"} do (including for HTTPS URLs).Server-side cached data is always accepted.Function \code{\link{download.file}} and contributed package\CRANpkg{RCurl} provide more comprehensive facilities to downloadfrom URLs.}\value{\code{file}, \code{pipe}, \code{fifo}, \code{url}, \code{gzfile},\code{bzfile}, \code{xzfile}, \code{unz} and \code{socketConnection}return a connection object which inherits from class\code{"connection"} and has a first more specific class.\code{open} and \code{flush} return \code{NULL}, invisibly.\code{close} returns either \code{NULL} or an integer status,invisibly. The status is from when the connection was last closed andis available only for some types of connections (e.g., pipes, files andfifos): typically zero values indicate success.\code{isOpen} returns a logical value, whether the connection iscurrently open.\code{isIncomplete} returns a logical value, whether the last readattempt was blocked, or for an output text connection whether there isunflushed output.}\section{Modes}{Possible values for the argument \code{open} are\describe{\item{\code{"r"} or \code{"rt"}}{Open for reading in text mode.}\item{\code{"w"} or \code{"wt"}}{Open for writing in text mode.}\item{\code{"a"} or \code{"at"}}{Open for appending in text mode.}\item{\code{"rb"}}{Open for reading in binary mode.}\item{\code{"wb"}}{Open for writing in binary mode.}\item{\code{"ab"}}{Open for appending in binary mode.}\item{\code{"r+"}, \code{"r+b"}}{Open for reading and writing.}\item{\code{"w+"}, \code{"w+b"}}{Open for reading and writing,truncating file initially.}\item{\code{"a+"}, \code{"a+b"}}{Open for reading and appending.}}Not all modes are applicable to all connections: for example URLs canonly be opened for reading. Only file and socket connections can beopened for both reading and writing. An unsupported mode is usuallysilently substituted.If a file or fifo is created on a Unix-alike, its permissions will bethe maximal allowed by the current setting of \code{umask} (see\code{\link{Sys.umask}}).For many connections there is little or no difference between text andbinary modes. For file-like connections on Windows, translation ofline endings (between LF and CRLF) is done in text mode only (but textread operations on connections such as \code{\link{readLines}},\code{\link{scan}} and \code{\link{source}} work for any form of lineending). Various \R operations are possible in only one of the modes:for example \code{\link{pushBack}} is text-oriented and is onlyallowed on connections open for reading in text mode, and binaryoperations such as \code{\link{readBin}}, \code{\link{load}} and\code{\link{save}} can only be done on binary-mode connections.The mode of a connection is determined when actually opened, which isdeferred if \code{open = ""} is given (the default for all but socketconnections). An explicit call to \code{open} can specify the mode,but otherwise the mode will be \code{"r"}. (\code{gzfile},\code{bzfile} and \code{xzfile} connections are exceptions, as thecompressed file always has to be opened in binary mode and noconversion of line-endings is done even on Windows, so the defaultmode is interpreted as \code{"rb"}.) Most operations that need writeaccess or text-only or binary-only mode will override the default modeof a non-yet-open connection.Append modes need to be considered carefully for compressed-fileconnections. They do \strong{not} produce a single compressed streamon the file, but rather append a new compressed stream to the file.Readers may or may not read beyond end of the first stream: currently\R does so for \code{gzfile}, \code{bzfile} and \code{xzfile}connections.}\section{Compression}{\R supports \command{gzip}, \command{bzip2} and \command{xz}compression (added in \R 2.10.0: also read-only support for itsprecursor \code{lzma} compression).For reading, the type of compression (if any) can be determined fromthe first few bytes of the file. Thus for \code{file(raw = FALSE)}connections, if \code{open} is \code{""}, \code{"r"} or \code{"rt"}the connection can read any of the compressed file types as well asuncompressed files. (Using \code{"rb"} will allow compressed files tobe read byte-by-byte.) Similarly, \code{gzfile} connections can readany of the forms of compression and uncompressed files in any readmode.(The type of compression is determined when the connection is createdif \code{open} is unspecified and a file of that name exists. If theintention is to open the connection to write a file with a\emph{different} form of compression under that name, specify\code{open = "w"} when the connection is created or\code{\link{unlink}} the file before creating the connection.)For write-mode connections, \code{compress} specifies how hard thecompressor works to minimize the file size, and higher values needmore CPU time and more working memory (up to ca 800Mb for\code{xzfile(compress = 9)}). For \code{xzfile} negative values of\code{compress} correspond to adding the \command{xz} argument\option{-e}: this takes more time (double?) to compress but mayachieve (slightly) better compression. The default (\code{6}) hasgood compression and modest (100Mb memory) usage: but if you are using\code{xz} compression you are probably looking for high compression.Choosing the type of compression involves tradeoffs: \command{gzip},\command{bzip2} and \command{xz} are successively less widely supported,need more resources for both compression and decompression, andachieve more compression (although individual files may buck thegeneral trend). Typical experience is that \code{bzip2} compressionis 15\% better on text files than \code{gzip} compression, and\code{xz} with maximal compression 30\% better. The experience with\R \code{\link{save}} files is similar, but on some large \file{.rda}files \code{xz} compression is much better than the other two. Withcurrent computers decompression times even with \code{compress = 9}are typically modest and reading compressed files is usually fasterthan uncompressed ones because of the reduction in disc activity.}\section{Encoding}{The encoding of the input/output stream of a connection can bespecified by name in the same way as it would be given to\code{\link{iconv}}: see that help page for how to find out whatencoding names are recognized on your platform. Additionally,\code{""} and \code{"native.enc"} both mean the \sQuote{native}encoding, that is the internal encoding of the current locale andhence no translation is done.Re-encoding only works for connections in text mode: reading from aconnection with re-encoding specified in binary mode will read thestream of bytes, but mixing text and binary mode reads (e.g., mixingcalls to \code{\link{readLines}} and \code{\link{readChar}}) is likelyto lead to incorrect results.The encodings \code{"UCS-2LE"} and \code{"UTF-16LE"} are treatedspecially, as they are appropriate values for Windows \sQuote{Unicode}text files. If the first two bytes are the Byte Order Mark\code{0xFEFF} then these are removed as some implementations of\code{\link{iconv}} do not accept BOMs. Note that whereas mostimplementations will handle BOMs using encoding \code{"UCS-2"} andchoose the appropriate byte order, some (including earlier versions of\code{glibc}) will not. There is a subtle distinction between\code{"UTF-16"} and \code{"UCS-2"} (see\url{https://en.wikipedia.org/wiki/UTF-16}: the use of characters inthe \sQuote{Supplementary Planes} which need surrogate pairs is veryrare so \code{"UCS-2LE"} is an appropriate first choice (as it is morewidely implemented).One caveat: \R's implementation of \code{"UCS-2LE"} and similar foroutput does not currently work on Windows, and on Unix it will defaultto Unix-style line endings. We recommend use of \code{UTF-8} instead.As from \R 3.0.0 the encoding \code{"UTF-8-BOM"} is accepted forreading and will remove a Byte Order Mark if present (which it oftenis for files and webpages generated by Microsoft applications). If aBOM is required (it is not recommended) when writing it should bewritten explicitly, e.g.\sspace{}by \code{writeChar("\ufeff", con, eos= NULL)} or \code{writeBin(as.raw(c(0xef, 0xbb, 0xbf)), binary_con)}Encoding names \code{"utf8"}, \code{"mac"} and \code{"macroman"} arenot portable, and not supported on all current \R platforms.\code{"UTF-8"} is portable and \code{"macintosh"} is the official(and most widely supported) name for \sQuote{Mac Roman}.Requesting a conversion that is not supported is an error, reportedwhen the connection is opened. Exactly what happens when therequested translation cannot be done for invalid input is in generalundocumented. On output the result is likely to be that up to theerror, with a warning. On input, it will most likely be all or someof the input up to the error.It may be possible to deduce the current native encoding from\code{\link{Sys.getlocale}("LC_CTYPE")}, but not all OSes record it.}\section{Blocking}{Whether or not the connection blocks can be specified for file, url(default yes), fifo and socket connections (default not).In blocking mode, functions using the connection do not return to the\R evaluator until the read/write is complete. In non-blocking mode,operations return as soon as possible, so on input they will returnwith whatever input is available (possibly none) and for output theywill return whether or not the write succeeded.The function \code{\link{readLines}} behaves differently in respect ofincomplete last lines in the two modes: see its help page.Even when a connection is in blocking mode, attempts are made toensure that it does not block the event loop and hence the operationof GUI parts of \R. These do not always succeed, and the whole \Rprocess will be blocked during a DNS lookup on Unix, for example.Most blocking operations on HTTP/FTP URLs and on sockets are subject to thetimeout set by \code{options("timeout")}. Note that this is a timeoutfor no response, not for the whole operation. The timeout is set atthe time the connection is opened (more precisely, when the lastconnection of that type -- \samp{http:}, \samp{ftp:} or socket -- wasopened).}\section{Fifos}{Fifos default to non-blocking. That follows S version 4 and isprobably most natural, but it does have some implications. Inparticular, opening a non-blocking fifo connection for writing (only)will fail unless some other process is reading on the fifo.Opening a fifo for both reading and writing (in any mode: one can onlyappend to fifos) connects both sides of the fifo to the \R process,and provides an similar facility to \code{file()}.}\section{Clipboard}{\code{file} can be used with \code{description = "clipboard"}#ifdef windowsin modes \code{"r"} and \code{"w"} only.#endif#ifdef unixin mode \code{"r"} only. This reads the X11 primary selection (see\url{http://standards.freedesktop.org/clipboards-spec/clipboards-latest.txt}),which can also be specified as \code{"X11_primary"} and the secondaryselection as \code{"X11_secondary"}. On most systems the clipboardselection (that used by \sQuote{Copy} from an \sQuote{Edit} menu) canbe specified as \code{"X11_clipboard"}.#endifWhen a clipboard is opened for reading, the contents are immediatelycopied to internal storage in the connection.#ifdef windowsWhen writing to the clipboard, the output is copied to the clipboardonly when the connection is closed or flushed. There is a 32Kb limiton the text to be written to the clipboard. This can be raised byusing e.g.\sspace{}\code{file("clipboard-128")} to give 128Kb.The clipboard works in Unicode wide characters, so encodings mightnot work as one might expect.#endif#ifdef unixUnix users wishing to \emph{write} to one of the X11 selections may beable to do so via \command{xclip}(\url{http://sourceforge.net/projects/xclip/}) or \command{xsel}(\url{http://www.vergenet.net/~conrad/software/xsel/}), for example by\code{pipe("xclip -i", "w")} for the primary selection.macOS users can use \code{pipe("pbpaste")} and\code{pipe("pbcopy", "w")} to read from and write to that system'sclipboard.#endif}\note{\R's connections are modelled on those in S version 4 (see Chambers,1998). However \R goes well beyond the S model, for example in outputtext connections and URL, compressed and socket connections.The default open mode in \R is \code{"r"} except for socket connections.This differs from S, where it is the equivalent of \code{"r+"},known as \code{"*"}.On (rare) platforms where \code{vsnprintf} does not return the neededlength of output there is a 100,000 byte output limit on the length ofa line for text output on \code{fifo}, \code{gzfile}, \code{bzfile} and\code{xzfile} connections: longer lines will be truncated with awarning.}\references{Chambers, J. M. (1998)\emph{Programming with Data. A Guide to the S Language.} Springer.Ripley, B. D. (2001) Connections. \emph{R News}, \bold{1/1}, 16--7.\url{https://www.r-project.org/doc/Rnews/Rnews_2001-1.pdf}}\seealso{\code{\link{textConnection}}, \code{\link{seek}},\code{\link{showConnections}}, \code{\link{pushBack}}.Functions making direct use of connections are (text-mode)\code{\link{readLines}}, \code{\link{writeLines}}, \code{\link{cat}},\code{\link{sink}}, \code{\link{scan}}, \code{\link{parse}},\code{\link{read.dcf}}, \code{\link{dput}}, \code{\link{dump}} and(binary-mode) \code{\link{readBin}}, \code{\link{readChar}},\code{\link{writeBin}}, \code{\link{writeChar}}, \code{\link{load}}and \code{\link{save}}.\code{\link{capabilities}} to see if \code{fifo} connections aresupported by this build of \R.\code{\link{gzcon}} to wrap \command{gzip} (de)compression around aconnection.\code{\link{options}} \code{HTTPUserAgent}, \code{internet.info} and\code{timeout} are used by some of the methods for URL connections.\code{\link{memCompress}} for more ways to (de)compress and referenceson data compression.#ifdef windowsTo flush output to the console, see \code{\link{flush.console}}.#endif}\examples{zz <- file("ex.data", "w") # open an output file connectioncat("TITLE extra line", "2 3 5 7", "", "11 13 17", file = zz, sep = "\n")cat("One more line\n", file = zz)close(zz)readLines("ex.data")unlink("ex.data")zz <- gzfile("ex.gz", "w") # compressed filecat("TITLE extra line", "2 3 5 7", "", "11 13 17", file = zz, sep = "\n")close(zz)readLines(zz <- gzfile("ex.gz"))close(zz)unlink("ex.gz")zz # an invalid connectionzz <- bzfile("ex.bz2", "w") # bzip2-ed filecat("TITLE extra line", "2 3 5 7", "", "11 13 17", file = zz, sep = "\n")close(zz)zz # print() method: invalid connectionprint(readLines(zz <- bzfile("ex.bz2")))close(zz)unlink("ex.bz2")## An example of a file open for reading and writingTfile <- file("test1", "w+")c(isOpen(Tfile, "r"), isOpen(Tfile, "w")) # both TRUEcat("abc\ndef\n", file = Tfile)readLines(Tfile)seek(Tfile, 0, rw = "r") # reset to beginningreadLines(Tfile)cat("ghi\n", file = Tfile)readLines(Tfile)Tfile # -> print() : "valid" connectionclose(Tfile)Tfile # -> print() : "invalid" connectionunlink("test1")## We can do the same thing with an anonymous file.Tfile <- file()cat("abc\ndef\n", file = Tfile)readLines(Tfile)close(Tfile)\dontrun{## fifo example -- may hang even with OS support for fifosif(capabilities("fifo")) {zz <- fifo("foo-fifo", "w+")writeLines("abc", zz)print(readLines(zz))close(zz)unlink("foo-fifo")}}#ifdef unix\donttest{## Unix examples of use of pipes# read listing of current directoryreadLines(pipe("ls -1"))# remove trailing commas. Suppose\dontshow{writeLines(c("450, 390, 467, 654, 30, 542, 334, 432, 421,","357, 497, 493, 550, 549, 467, 575, 578, 342,","446, 547, 534, 495, 979, 479"), "data2_")}\dontrun{\% cat data2_450, 390, 467, 654, 30, 542, 334, 432, 421,357, 497, 493, 550, 549, 467, 575, 578, 342,446, 547, 534, 495, 979, 479}# Then read this byscan(pipe("sed -e s/,$// data2_"), sep = ",")\dontshow{unlink("data2_")}# convert decimal point to comma in output: see also write.table# both R strings and (probably) the shell need \ doubledzz <- pipe(paste("sed s/\\\\\\\\./,/ >", "outfile"), "w")cat(format(round(stats::rnorm(48), 4)), fill = 70, file = zz)close(zz)file.show("outfile", delete.file = TRUE)}\dontrun{## example for a machine running a finger daemoncon <- socketConnection(port = 79, blocking = TRUE)writeLines(paste0(system("whoami", intern = TRUE), "\r"), con)gsub(" *$", "", readLines(con))close(con)}#endif\dontrun{## Two R processes communicating via non-blocking sockets# R process 1con1 <- socketConnection(port = 6011, server = TRUE)writeLines(LETTERS, con1)close(con1)# R process 2con2 <- socketConnection(Sys.info()["nodename"], port = 6011)# as non-blocking, may need to loop for inputreadLines(con2)while(isIncomplete(con2)) {Sys.sleep(1)z <- readLines(con2)if(length(z)) print(z)}close(con2)## examples of use of encodings# write a file in UTF-8cat(x, file = (con <- file("foo", "w", encoding = "UTF-8"))); close(con)# read a 'Windows Unicode' fileA <- read.table(con <- file("students", encoding = "UCS-2LE")); close(con)}}\keyword{file}\keyword{connection}