Probability-proportional-to-size ranked-set sampling from stratified populations
Section 1. Introduction
In
survey sampling studies, selection of a sampling design depends on the
structure of the population. In this paper, we consider a population structure
having two main features. It must contain a size variable
and a variable of interest
The values of the size variable should be
approximately proportional to the values of the
-variable, and the values of
the
-variable should be available
for all population units prior to sampling. The second feature of the
population structure is that a small percentage of population units should
produce extreme values in both the
- and
-variables with different
proportionality constants. These units usually produce larger means and
variances than the rest of the units in the population in both the
- and
-variables. This population
structure is very common in practice. In agricultural sampling, a farm
population in a state or country may contain two variables, the crop production
and the farm size
in acres. Farms can be divided into two
groups, the farms that have small/normal sizes and the mega-farms that have
extremely large values in the
- and
-variables. The percentage of
the mega-farms would be small, but they may have larger means and variances in
the
- and
-variables and the
proportionality constant between the
- and
-values may be larger.
The
population structure in the Monthly Retail Trade Survey performed by the United
States Census Bureau would be another example. In this case, the population is
defined by the business establishments that have Employer Identification
Numbers (EINs). The Census Bureau uses a very complex design in which the
previous years’ annual revenues are used to construct a size variable. The
structure of the population fits the setting we consider in this paper.
Revenues from the previous years would be approximately proportional to the
current revenues. The revenues for most of the businesses would take typical
values, while revenues of a certain percentage of businesses would be extremely
large, producing a larger mean and variance and a different proportionality
constant. Different proportionality constant may happen because large
businesses may be more/less productive than the rest of the businesses.
We
provide a third example using apple production data in Turkey in 2002. The data
set was collected by the Turkish Statistical Institute and reported in Kadilar
and Cingi (2003) and Ozturk and Bayramoglu-Kavlak (2018). It contains two
variables, apple production
(in 1,000kg) and the numbers of apple trees
in townships (or localities). The sampling
units are the 851 localities in the data set. The
-values in all townships are
available in the sampling frame prior to sampling. Figure 1.1 provides the
scatter plot of the
- and
-values, where we see that the
value of the
-variable is an increasing
function of the
-value. We also observe that
red-colored points marked with
in the plot have large values for both
- and
-variables and their
proportionality constant is different from the other points. Hence, this
population fits into our population structure.
The
apple production data have additional structures. The entire population is
stratified into seven different geographical regions: Marmara, Aegean,
Mediterranean, Central Anatolia, Black Sea, Eastern Anatolia and Southeastern
Anatolia. These regions have different climate patterns and apple production
changes significantly from one region to another. The extreme observations
marked with
in Figure 1.1 come from the Marmara,
Aegean, Mediterranean and Central Anatolia regions. This is a natural setting
to construct a PPS-ranked-set sample from each sub-population.

Description for Figure 1.1
Scatter
plot of the apple production. The apple production in 1,000kg is on the y-axis,
ranging from 0 to 300,000. The number of apple trees is on the x-axis, ranging
from 0 to 3,000,000. The graph shows that the value of the apple production is
an increasing function of the number of apple trees. There are extreme values
indicated by red-colored points marked with
in the plot that have large values for both
- and
-variables.
They come from the Marmara, Aegean, Mediterranean and Central Anatolia regions.
The variable
in our population setting provides information
about the relative size of the units in the entire population. Since the
variable
is approximately proportional to the variable
the size of the unit indicates the importance
of its contribution to the variable of interest
Hence, important units should have higher
probability of being included in the sample. Probability-proportional-to-size
(PPS) sampling deliberately imposes higher selection probabilities for
important units to produce unbiased and highly efficient estimators for the
population mean and/or total. The main contribution of this paper is to
introduce a new sampling design which combines the ranking information in a
ranked-set sample (RSS) with the advantage of unequal selection probabilities
as used in a PPS sample.
A typical PPS sample contains
the triplets
where
and
are the value of
and the selection probability of the unit
for each draw under sampling with replacement
selection. Though PPS sampling can be done without replacement, here all
references to PPS sampling refers to sampling with replacement. Readers may
refer to Thompson (2002, page 53) for the details of PPS sampling. In a
PPS sample, the size variable is not necessarily used directly in the
construction of the estimators. On the other hand, the values of the
-variable are available for
all population units even for the units not included in the PPS sample. Hence
the
-variable could help us to
borrow additional information from a comparison set of
unmeasured units.
For the construction of a
typical data point
in a PPS-ranked-set sample, we select
units from the population using PPS sampling
with replacement to form a comparison set
We rank these units without measurement based
on the
-variable with no additional
cost. We then measure the value of the
-variable for only one unit,
the unit having rank
obtaining
where
and
are the value of the
-variable and the selection
probability of the unit
at each draw. A data point in the comparison
set,
provides more information than a data point in
a PPS sample,
since the rank
borrows additional information from the other
unmeasured units in the comparison set. In
this paper, we use this idea to construct a PPS-ranked-set sample that is more
informative than a PPS sample. The details of this sampling procedure will be
provided in Section 2.
The position information is
used in a slightly different context in ranked-set sample (RSS) and
judgment-post-stratified (JPS) sampling designs to borrow information from the
unmeasured population units. Construction of a ranked-set sample of size
requires one to determine two integers
and
where
and
are the set and cycle sizes, respectively. The
set size
controls the amount of information that can be
borrowed from the units in comparison sets. The cycle size
is used to increase the total sample size in a
RSS. Once
and
are chosen, one then selects
units from the population and partitions them
into
disjoint comparison sets, each having
units. Units in each comparison set are ranked
without measurement using the
-variable and the value of the
-variable
associated with the
ranked
is measured in
different comparison sets,
The measured values
are called a ranked-set sample.
The construction of a JPS
sample of size
starts with a simple random sample of size
and measures all of them,
For each measured unit in this sample, one
then selects an additional
units from the population to form a comparison
set of size
The rank
of the measured unit
in each of these comparison sets is
determined. The pairs of
constitute a JPS sample.
The
RSS and JPS samples create induced order statistics for the
-variable through the ranks of
the
-variable in comparison sets.
Hence, the random variable
given that
for the JPS sample) is stochastically smaller
than the random variable
given that
for
This stochastic ordering property induces an
implicit stratification among the measured sample units. Efficiency
improvements of RSS and JPS samples over a simple random sample can be
anticipated from the partition of the total variation into between- and
within-strata variation. For further details on these sampling designs, readers
may refer to the review paper in Wolfe (2012) and references therein. Both RSS
and JPS samples use the position information of the units in the comparison
sets, but they do not completely use the information provided by the selection
probabilities in a PPS sample. All units in comparison sets for RSS and JPS are
selected with equal probabilities. Hence, they may not be appropriate for the
population structure that we consider in this paper.
MacEachern,
Stasny and Wolfe (2004) introduced the JPS design in an infinite population
setting. In a finite population setting, constructions of the JPS and RSS
samples depend on whether the comparison sets are selected with or without
replacement. Patil, Sinha and Taillie (1995) considered an RSS design in a
finite population, where none of the units in a comparison set is returned to
the population prior to selection of the next comparison set. Deshpande, Frey
and Ozturk (2006) expanded the RSS sampling design with three different without
replacement selection policies and constructed nonparametric confidence
intervals for population quantiles.
Probability
sampling has also generated extensive research interest in RSS and JPS
sampling. Al-Saleh and Samawi (2007), Ozdemir and Gokpinar (2007 and 2008),
Gokpinar and Ozdemir (2010), Ozturk and Jafari Jozani (2013), Frey (2011)
and Ozturk (2014) computed inclusion probabilities and constructed
Horwitz-Thompson type estimators for the population mean and total based on a
ranked-set sample. These research papers show that an RSS design yields a
substantial amount of improvement in efficiency over the usual simple random
sampling design. Ozturk (2016) developed estimators for the population mean based
on a JPS sample, where he showed that the estimator needs a finite population
correction factor similar to the one used in a simple random sample.
A
few researchers have applied the RSS methodology to existing survey sampling
designs. Muttlak and McDonald (1992) incorporated the RSS sampling design with
a line intersect method. Sroka (2008) used it in stratified sampling by
constructing an RSS sample from each stratum. Wang, Lim and Stokes (2016)
considered the RSS design in a cluster randomized design with a mixed effect
model, where the cluster effect is treated as random. They showed that use of
RSS at the cluster level has much bigger impact on efficiency than using the
RSS at the within-cluster level. Nematollahi, Salehi and Aliakbari Saba
(2008) used the RSS design in a finite population setting only in the second
stage of a two-stage sampling with replacement selection scheme. Since they use
the RSS design only in the second stage with replacement, the efficiency
improvement of their estimator with respect to a two-stage SRS sample estimator
was minimal. Sud and Mishra (2006) also used a two-stage cluster sample with
ranked set sampling design in a finite population setting under the assumption
that the cluster population sizes are all equal. Ozturk (2019a) developed
design based statistical inference for a two-stage clustered ranked-set sample
in a finite population setting.
In
this paper, we develop statistical inference for the PPS-ranked-set sampling
design in a population setting where the values of the size variable are
roughly proportional to the values of the variable of interest and a small
percentage of population units produces large
- and
-values with a different
proportionality constant. We motivate the new sampling design using apple
production data. Section 2 introduces the PPS-ranked-set sample in a
finite population setting. It constructs unbiased estimators for the population
mean, total and their variances. We show that the PPS-ranked-set sample
estimator has smaller variance than a PPS sample estimator. Section 3 extends the PPS-ranked-set sample to a stratified population and constructs
unbiased estimators for the population mean, total and their variances.
Section 4 considers four different sample size allocation procedures to
minimize the variance of the estimator under a cost model and different stratum
population structures. Section 5 provides an efficiency comparison for the
PPS-ranked-set sample estimator of the population mean with respect to other
competing estimators. Section 6 illustrates the use of PPS-ranked-set
sample data to estimate the apple production in Turkey. Section 7 provides
some concluding remarks. All proofs are given in the Appendix.
ISSN : 1492-0921
Editorial policy
Survey Methodology publishes articles dealing with various aspects of statistical development relevant to a statistical agency, such as design issues in the context of practical constraints, use of different data sources and collection techniques, total survey error, survey evaluation, research in survey methodology, time series analysis, seasonal adjustment, demographic studies, data integration, estimation and data analysis methods, and general survey systems development. The emphasis is placed on the development and evaluation of specific methodologies as applied to data collection or the data themselves. All papers will be refereed. However, the authors retain full responsibility for the contents of their papers and opinions expressed are not necessarily those of the Editorial Board or of Statistics Canada.
Submission of Manuscripts
Survey Methodology is published twice a year in electronic format. Authors are invited to submit their articles in English or French in electronic form, preferably in Word to the Editor, (statcan.smj-rte.statcan@canada.ca, Statistics Canada, 150 Tunney’s Pasture Driveway, Ottawa, Ontario, Canada, K1A 0T6). For formatting instructions, please see the guidelines provided in the journal and on the web site (www.statcan.gc.ca/SurveyMethodology).
Note of appreciation
Canada owes the success of its statistical system to a long-standing partnership between Statistics Canada, the citizens of Canada, its businesses, governments and other institutions. Accurate and timely statistical information could not be produced without their continued co-operation and goodwill.
Standards of service to the public
Statistics Canada is committed to serving its clients in a prompt, reliable and courteous manner. To this end, the Agency has developed standards of service which its employees observe in serving its clients.
Copyright
Published by authority of the Minister responsible for Statistics Canada.
© Her Majesty the Queen in Right of Canada as represented by the Minister of Industry, 2020
Use of this publication is governed by the Statistics Canada Open Licence Agreement.
Catalogue No. 12-001-X
Frequency: Semi-annual
Ottawa