Learn R Programming

tm (version 0.7-17)

Text Mining Package

Description

A framework for text mining applications within R.

Copy Link

Version

Install

install.packages('tm')

Monthly Downloads

60,335

Version

0.7-17

License

GPL-3

Maintainer

Kurt Hornik

Last Published

December 10th, 2025

Functions in tm (0.7-17)

Corpus

Corpora
PlainTextDocument

Plain Text Documents
PCorpus

Permanent Corpora
DirSource

Directory Source
Source

Sources
SimpleCorpus

Simple Corpora
DataframeSource

Data Frame Source
Reader

Readers
URISource

Uniform Resource Identifier Source
Docs

Access Document IDs and Terms
Zipf_n_Heaps

Explore Corpus Term Frequency Characteristics
VectorSource

Vector Source
VCorpus

Volatile Corpora
acq

50 Exemplary News Articles from the Reuters-21578 Data Set of Topic acq
ZipSource

ZIP File Source
TextDocument

Text Documents
XMLSource

XML Source
WeightFunction

Weighting Function
XMLTextDocument

XML Text Documents
findMostFreqTerms

Find Most Frequent Terms
foreign

Read Document-Term Matrices
content_transformer

Content Transformers
crude

20 Exemplary News Articles from the Reuters-21578 Data Set of Topic crude
inspect

Inspect Objects
hpc

Parallelized ‘lapply’
tm_combine

Combine Corpora, Documents, Term-Document Matrices, and Term Frequency Vectors
getTokenizers

Tokenizers
getTransformations

Transformations
plot

Visualize a Term-Document Matrix
findAssocs

Find Associations in a Term-Document Matrix
findFreqTerms

Find Frequent Terms
readDOC

Read In a MS Word Document
stemCompletion

Complete Stems
removeWords

Remove Words from a Text Document
readTagged

Read In a POS-Tagged Word Text Document
stripWhitespace

Strip Whitespace from a Text Document
readReut21578XML

Read In a Reuters-21578 XML Document
readPlain

Read In a Text Document
termFreq

Term Frequency Vector
readRCV1

Read In a Reuters Corpus Volume 1 Document
removeSparseTerms

Remove Sparse Terms from a Term-Document Matrix
writeCorpus

Write a Corpus to Disk
TermDocumentMatrix

Term-Document Matrix
weightTfIdf

Weight by Term Frequency - Inverse Document Frequency
removePunctuation

Remove Punctuation Marks from a Text Document
meta

Metadata Management
stopwords

Stopwords
stemDocument

Stem Words
tm_reduce

Combine Transformations
weightTf

Weight by Term Frequency
weightSMART

SMART Weightings
tm_term_score

Compute Score for Matching Terms
weightBin

Weight Binary
tokenizer

Tokenizers
readPDF

Read In a PDF Document
readDataframe

Read In a Text Document from a Data Frame
readXML

Read In an XML Document
tm_filter

Filter and Index Functions on Corpora
removeNumbers

Remove Numbers from a Text Document
tm_map

Transformations on Corpora