Learn R Programming

sfa (version 1.2.0)

data_gen_cs: Generate Cross-Sectional Data for Stochastic Frontier Analysis

Description

data_gen_cs generates simulated cross-sectional data based on the stochastic frontier model, allowing for different distributional assumptions for the one-sided technical inefficiency error term (\(u\)) and the two-sided idiosyncratic error term (\(v\)). The model has the general form: \(Y = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + v - u\) where \(u \geq 0\) and represents inefficiency. All variants are produced so that the user can select those that they want.

Usage

data_gen_cs(N, rand, sig_u, sig_v, cons, beta1, beta2, a, mu, sig_w = sig_u,
            shape_g = 2, m_nak = 1, mu_ln = -0.5, k_w = 1.5, lam_tsl = 1.5)

Value

A data frame containing \(N\) observations with the following columns:

name

Individual identifier (simply \(1\) to \(N\)).

cons

The constant term value.

x1

Simulated explanatory variable \(x_1\).

x2

Simulated explanatory variable \(x_2\).

u, uz, u_t, u_c, u_e, u_u, u_tn, u_st, u_thn

The simulated one-sided error terms under different distributions.

u_r

Rayleigh one-sided error term, for y_pcs_r. Drawn on the convention sfm's "NR" uses, \(E[u^2] = \sigma_u^2\), so the Rayleigh scale is \(\sigma_u/\sqrt{2}\), \(E[u] = \sigma_u\sqrt{\pi}/2\) and \(Var(u) = (1-\pi/4)\sigma_u^2\).

v, v_t, v_c, v_st, v_thn

The simulated two-sided error terms under different distributions.

lam_st

The common Gamma(a/2, a/2) scale-mixing variable behind y_pcs_st; both of that column's components are divided by its square root, which is what makes the composed error skew-t.

y_pcs, y_pcs_t, y_pcs_st, y_pcs_thn, y_pcs_e, y_pcs_ez, y_pcs_c, y_pcs_u, y_pcs_z, y_pcs_w, y_pcs_tn

The dependent variable \(Y\) under the corresponding SFA model distributions.

y_pcs_r

Normal-Rayleigh dependent variable, for sfm's "NR". Use this column, not y_pcs: the normal-Rayleigh and normal-half-normal composed errors are different families, not reparameterizations of one another, and "NR" cannot recover y_pcs's parameters. True values are \((\sigma_v, \sigma_u) = \) (sig_v, sig_u).

z

The auxiliary variable used for heteroskedasticity in y_pcs_z, y_pcs_ez, y_tthn_z, and y_zisf_z.

uz_e

Exponential one-sided error term with heteroskedastic scale, for y_pcs_ez.

w_tt, w_tt_hn, wz_hn

The second one-sided error component (\(w\)) for the two-tier columns below, under exponential, homoskedastic half-normal, and heteroskedastic half-normal distributions respectively.

zp

A second auxiliary variable used for heteroskedasticity of \(\sigma_w\) in y_tthn_z.

y_ttne

Homoskedastic two-tier dependent variable (\(v + w - u\), both one-sided components exponential), for ttsfm's "TTNE"/"TTNLS".

y_tthn

Homoskedastic two-tier dependent variable (both one-sided components half-normal), for ttsfm's "TTHN" with no pipes.

y_tthn_z

Heteroskedastic two-tier dependent variable (\(\sigma_u\) a function of z, \(\sigma_w\) a function of zp), for ttsfm's "TTHN" with formula ~x1+x2|z|zp.

eff_ind, eff_ind_z

Indicator (1 = efficient, \(u\) forced to 0) used to build the zero-inefficiency columns below.

prob_z_true

The true heteroskedastic "efficient regime" probability used to draw eff_ind_z.

y_zisf

Zero-inefficiency dependent variable with constant efficient-regime probability, for zsfm's "ZISF".

y_zisf_z

Zero-inefficiency dependent variable with z-dependent efficient-regime probability, for zsfm's "ZISF_Z".

u_g, y_pcs_g

Gamma one-sided error term and the corresponding dependent variable, for sfm's "NG". Shape shape_g, scale sig_u/shape_g, so \(E[u] = \sigma_u\) whatever shape is chosen.

u_nak, y_pcs_nak

Nakagami one-sided error term and dependent variable, for sfm's "NNAK". Shape m_nak, spread \(\Omega = \sigma_u^2\). The default m_nak = 1 is the Rayleigh case, so this column coincides in distribution with y_pcs_r at the default.

u_ge, y_pcs_ge

Generalized-exponential one-sided error term and dependent variable, for sfm's "NGE", drawn as \(u = -\sigma_u\log(1 - \sqrt{U})\) with \(U \sim \mathrm{Unif}(0,1)\).

u_ln, y_pcs_ln

Lognormal one-sided error term and dependent variable, for sfm's "NLN". Log-scale mean mu_ln and log-scale standard deviation sig_u.

u_w, y_pcs_wb

Weibull one-sided error term and dependent variable, for sfm's "NW". Shape k_w, scale sig_u.

con

A constant column set to 1, potentially for use in estimation.

Arguments

N

A single integer specifying the number of observations (cross-sectional units).

rand

A single integer to set the seed for the random number generator, ensuring reproducibility.

sig_u

The standard deviation parameter (\(\sigma_u\)) for the base distribution of the one-sided error term \(u\).

sig_v

The standard deviation parameter (\(\sigma_v\)) for the base distribution of the two-sided error term \(v\).

cons

The value of the constant term (intercept) in the model.

beta1

The coefficient for the \(x_1\) variable.

beta2

The coefficient for the \(x_2\) variable.

a

The degrees of freedom parameter for the t half-t distribution (u_t and v_t, respectively). Requires the rt function.

mu

The mean parameter (\(\mu\)) for the normal truncated normal distribution (u_tn). Requires the rtruncnorm function.

sig_w

The standard deviation/scale parameter (\(\sigma_w\)) for the second one-sided error component \(w\), used to generate two-tier model columns (see ttsfm). Defaults to sig_u.

shape_g

Shape of the gamma inefficiency draw behind y_pcs_g, which targets sfm(model_name = "NG"). The scale is set to sig_u/shape_g, so \(E[u] = \sigma_u\) regardless of the shape chosen. sfm()'s NG likelihood reports the shape as mu and the scale as sigu, so the values to recover are shape_g and sig_u/shape_g.

m_nak

Nakagami shape \(m\) behind y_pcs_nak, which targets sfm(model_name = "NNAK"). The spread is \(\Omega = \sigma_u^2\), so \(\sigma_u\) is the root-mean-square of \(u\). The default 1 is the Rayleigh case; 0.5 would collapse the Nakagami onto the half-normal and so duplicate y_pcs.

mu_ln

Log-scale mean of the lognormal inefficiency draw behind y_pcs_ln, which targets sfm(model_name = "NLN"). The log-scale standard deviation is sig_u. The default -0.5 with sig_u = 1 gives \(E[u] = 1\).

k_w

Weibull shape behind y_pcs_wb, which targets sfm(model_name = "NW"). The scale is sig_u.

lam_tsl

Skew parameter \(\lambda\) behind y_pcs_tsl, which targets sfm(model_name = "TSL"). The scale is sig_u; \(\lambda \to 0\) recovers the exponential case.

Author

David H. Bernstein

Details

The function simulates two explanatory variables, \(x_1\) and \(x_2\), as transformations of uniform random variables.

The function generates several different frontier models by combining various distributions for \(u\) and \(v\):

  • **\(u\) Distributions (Inefficiency):** Half-Normal (HN), Truncated Normal (TN), Half-T (HT), Half-Cauchy (HC), Exponential (E), Half-Uniform (HU).

  • **\(v\) Distributions (Idiosyncratic):** Normal (N), t, Cauchy (C).

**Specific Model Outputs (y_pcs variants):**

  • y_pcs: Normal-Half Normal (N-HN): \(v \sim N(0, \sigma_v^2)\), \(u \sim |N(0, \sigma_u^2)|\).

  • y_pcs_z: N-HN with Heteroskedastic \(\sigma_u\): \(\sigma_{u,i} = \exp(0.9 + 0.6 Z_i)\), where \(Z\) is a uniform variable.

  • y_pcs_t: T-Half T (T-HT): \(v \sim T(\text{df}=a) \cdot \sigma_v\), \(u \sim |T(\text{df}=a)| \cdot \sigma_u\).

  • y_pcs_st: the column that actually matches sfm's THT. y_pcs_t above draws its two components as two independent rt() variates, which shares the degrees of freedom but not the mixing variable, so the composed error is not skew-t and THT cannot recover it. y_pcs_st uses the single common \(\lambda \sim \Gamma(a/2, a/2)\) of Tancredi (2002), i.e. \(y = h + (\epsilon - z)/\sqrt{\lambda}\). y_pcs_t is retained unchanged only for backward compatibility.

  • y_pcs_thn: Student's t--half normal, matching sfm's tHN: \(v \sim T(\text{df}=a) \cdot \sigma_v\) and \(u \sim |N(0,\sigma_u^2)|\), drawn independently. Heavy-tailed noise with a conventional one-sided term -- distinct from both y_pcs_t and y_pcs_st, in which the inefficiency is heavy-tailed too.

  • y_pcs_tn: Normal-Truncated Normal (N-TN): \(v \sim N(0, \sigma_v^2)\), \(u \sim TN(\mu, \sigma_u^2)\) on \([0, \infty)\).

  • y_pcs_e: Normal-Exponential (N-E): \(v \sim N(0, \sigma_v^2)\), \(u \sim Exp(\phi)\), where \(\phi = 1/\sigma_u\).

  • y_pcs_c: Cauchy-Half Cauchy (C-HC): \(v \sim Cauchy(0, \sigma_v)\), \(u \sim |Cauchy(0, \sigma_u)|\).

  • y_pcs_u: Normal-Half Uniform (N-HU): \(v \sim N(0, \sigma_v^2)\), \(u \sim U(0, \sigma_u)\).

  • y_pcs_w: Normal + Cauchy - Half Normal: \(v \sim N(0, \sigma_v^2) + Cauchy(0, \sigma_v)\), \(u \sim |N(0, \sigma_u^2)|\). This introduces a composite \(v\) term.

**Note:** The rtruncnorm function is required for y_pcs_tn and loads with the package. In isolation it could be loaded by using library(truncnorm).

See Also

data_gen_p for the panel generator, and sfm, zsfm and ttsfm for the estimators these columns are built for.

rnorm, runif, rt, rexp, rcauchy, rtruncnorm (if available).

Examples

Run this code

# Generate 100 observations of SFA data
data_sfa <- data_gen_cs(
  N     = 100,
  rand  = 123,
  sig_u = 0.5,
  sig_v = 0.2,
  cons  = 5,
  beta1 = 1.5,
  beta2 = 2.0,
  a     = 5,   # degrees of freedom for T/Half-T
  mu    = 0.1  # mean for Truncated Normal
)

# Display the first few rows of the generated data
head(data_sfa)

# Example of a Normal-Half Normal SFA model data
summary(data_sfa$y_pcs)
plot(density(data_sfa$y_pcs))

Run the code above in your browser using DataLab