American Educational Research Association, American Psychological Association, National Council on Measurement in Education [AERA, APA, NCME]. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
BakerE. L.BartonP. E.Darling-HammondL.HaertelE.LaddH. F.LinnR. L.…ShepardL. A. (2010). Problems with the use of student test scores to evaluate teachers. Retrieved February 4, 2015, from the Economic Policy Institutehttp://www.epi.org/publication/bp278/
4.
BallouD.SpringerM. G. (2015). Using student test scores to measure teacher performance: Some problems in the design and implementation of evaluation systems. Educational Researcher, 44, 77–86.
5.
BeatonA. E.LinnR. L.BohrnstedtG. W. (2012). Alternative approaches to setting performance standards for the National Assessment of Educational Progress (NAEP). Washington, DC: American Institutes for Research, NAEP Validity Studies Panel.
6.
BockR. D.ThissenD.ZimowskiM. F. (1997). IRT estimation of domain scores. Journal of Educational Measurement, 34, 197–211.
7.
BraunH. (2015). The value in value added depends on the ecology. Educational Researcher, 44, 127–131.
8.
BriggsD. C. (2013). Measuring growth with vertical scales. Journal of Educational Measurement, 50, 204–226.
9.
BriggsD. C. (2015, 4). Debate: Equal interval scales in educational testing: Attainable goal or myth?Presentation at the annual meeting of the National Council on Measurement in Education, Chicago, IL.
10.
BriggsD. C.DomingueB. (2013). The gains from vertical scaling. Journal of Educational and Behavioral Statistics, 38, 551–576.
11.
CastellanoK. E.HoA. D. (2015). Practical differences among aggregate-level conditional status metrics: From median student growth percentiles to value-added models. Journal of Educational and Behavioral Statistics, 40, 35–68.
12.
Darling-HammondL. (2015). Can value added add value to teacher evaluation?Educational Researcher, 44, 132–137.
13.
GoldhaberD. (2015). Exploring the potential of value-added performance measures to affect the quality of the teacher workforce. Educational Researcher, 44, 87–95.
14.
GoldringE.GrissomJ. A.RubinM.NeumerskiC. M.CannataM.DrakeT.SchuermannP. (2015). Make room value added: Principals’ human capital decisions and the emergence of teacher observation data. Educational Researcher, 44, 96–104.
15.
HabermanS. J. (2008). When can subscores have value?Journal of Educational and Behavioral Statistics, 33, 204–229.
16.
HarrisD. N.HerringtonC. D. (2015). Editors’ introduction: The use of teacher value- added measures in schools: New evidence, unanswered questions, and future prospects. Educational Researcher, 44, 71–76.
JohnsonS. M. (2015). Will VAMS reinforce the walls of the egg-crate school?Educational Researcher, 44, 117–126.
19.
LawleyD. N. (1943). On problems connected with item selection and test construction. Proceedings of the Royal Society of Edinburgh, 62-A, 74–82.
20.
LazarsfeldP. F. (1950). The logical and mathematical foundation of latent structure analysis. In StoufferS. A.GuttmanL.SuchmanE. A.LazarsfeldP. F.StarS. A.ClausenJ. A., Measurement and prediction (pp. 362–412). New York, NY: Wiley.
21.
LordF. M.NovickM. R. (1968). Statistical theories of mental test scores. Reading, MA: Addison-Wesley.
22.
Maydeu-OlivaresA. (2015). Evaluating the fit of IRT models. In ReiseS. P.RevickiD. A. (Eds.), Handbook of item response theory modeling: Applications to typical performance assessment (pp. 111–127). New York, NY: Taylor & Francis (Routledge).
23.
McCaffreyD. F.LockwoodJ. R.KoretzD. M.HamiltonL. S. (2003). Evaluating value-added models for teacher accountability. Retrieved February 4, 2015, from The RAND Corporationhttp://www.rand.org/pubs/monographs/2004/RAND_MG158.pdf
24.
National Research Council. (2005). Measuring literacy: Performance levels for adults. In Committee on Performance Levels for Adult Literacy, HauserR. M.EdleyC. F.JrKoenigJ. A.ElliottS. W. (Eds.), Board on Testing and Assessment, Center for Education, Division of Behavioral and Social Sciences and Education. Washington, DC: The National Academies Press.
25.
National Research Council. (2011). Incentives and test-based account-ability in Education. In HoutM.ElliottS. W. (Eds.), Board on Testing and Assessment, Division of Behavioral and Social Sciences and Education. Washington, DC: The National Academies Press.
26.
PommerichM.NicewanderW. A.HansonB. A. (1999). Estimating average domain scores. Journal of Educational Measurement, 36, 199–216.
27.
QuinnH. O. (2014). Bifactor models, explained common variance (ECV), and the usefulness of scores from unidimensional item response theory analyses. Unpublished Master’s thesis, The University of North Carolina at Chapel Hill, Chapel Hill, NC.
28.
RaudenbushS. W. (2015). Value added: A case study in the mismatch between education research and policy. Educational Researcher, 44, 138–141.
29.
ReiseS. P.MooreT. M.HavilandM. G. (2010). Bifactor models and rotations: Exploring the extent to which multidimensional data yield univocal scale scores. Journal of Personality Assessment, 92, 544–559.
30.
SymondsP. M. (1929). Choice of items for a test on the basis of difficulty. Journal of Educational Psychology, 20, 481–493.
31.
ThorndikeE. L. (1918). The nature, purposes, and general methods of measurements of educational products. In WhippleG. M. (Ed.), The Seventeenth yearbook of the National Society for Study of Education. Part II. The measurement of educational products (pp. 16–24). Bloomington, IL: Public School.
32.
ThurstoneL. L. (1925). A method of scaling psychological and educational tests. Journal of Educational Psychology, 16, 433–449.
33.
van der LindenW. (2015, 4). Debate: Equal interval scales in educational testing: Attainable goal or myth?Presentation at the annual meeting of the National Council on Measurement in Education, Chicago, IL.