Pondering the Usefulness of Value-Added Assessment of Teachers

Value-added teacher assessment has been a mantra for education “reformers” throughout the debate over Race to the Top. We’ve got to evaluate teachers and make hiring and firing decisions on the basis of real student performance measures – you know, like businesses – like the real world does! (A highly questionable assumption indeed – AIG bonuses anyone?).

I address the technical issues with value-added assessment of teachers here, indicating just how premature these assertions are from a technical standpoint.

https://schoolfinance101.wordpress.com/2009/11/07/teacher-evaluation-with-value-added-measures/

At present, good value added measures are little more than  a really cool (if not totally awesome) research tool, but most of the best analyses of value-added as a tool for teacher evaluation suggest that even in the best of cases there still exist potentially problematic biases.

Let’s set these technical issues aside for now and explore some practical issues. For example, just how many teachers in a public education system could even be evaluated with value-added assessment? Consider these constraints.

  1. Most states, like New Jersey, implement yearly assessments in grades 3 through 8, and perhaps end of course or some HS exit exam. (I’ll set aside concerns over the fact that annual, rather than fall-spring assessment captures vast differences in summer learning which play out by student economic status – advantaging some teachers and disadvantaging others, depending on which kids they have).
  2. In most cases, the established and more reliable tests exist only in language arts and math, though some states have implemented science and/or social studies tests which are arguably less cumulative.
  3. The most reliable VA assessment of teachers occurs where there exist multiple points of historical scores on students prior to the observed teacher  (smaller technical point). This really casts doubt on the usefulness of VA assessment for evaluating teachers who have kids in their first few years of being assessed (grades 3 and 4 in NJ and many states).
  4. by the time a student hits middle school, they typically interact with multiple teachers who may have simultaneous influences on each others’ content area success. Even if we ignore this, at best we can look at the language arts and math teachers in the middle school setting.
  5. you have to jump over those untested grade 9 and 10 students and their teachers. If we have end of course exams, we don’t know what the beginning of course status necessarily was – at least in a VA modeling sense.

So, here is a listing of the certified staffing in New Jersey (below) in 2008 based on their grade levels and areas of teaching. The list does not include everyone, but does capture the main assignment (JOB Code 1) for the vast majority of school assigned teaching (and principal) personnel.

What this list shows us is that in the best possible case, in a state with annual Grades 3 to 8 assessment and shifting to end of course exams, we might be able to generate VA estimates of effectiveness for about 10% or 20% (just saw that “ungraded elementary” group) of the teachers. That is, 10% (up to 20%) would be subject to a different evaluation system than the rest. In fact, nearly 50% of teachers would be infeasible to evaluate at all. Indeed they are an important 10% (or perhaps 20%).

Okay, so maybe this would create incentive for the real gunners in the mix of potential teachers to dive into those areas evaluated by VA. There exists an equal if not stronger possibility that the real gunners in the mix of potential teachers will avoid those classrooms of kids, schools or districts where – in the evaluated content areas and grade levels – they face an uphill battle to improve outcomes (hopefully, some will welcome the challenge).

There are some obvious solutions to this dilemma –

  1. Test everything, every year by cumulative measures, fall and spring. Okay. That seems a bit absurd, but it might be a good economic stimulus for the testing industry. I still struggle with how we would evaluate teachers in supporting roles, as many of those listed below or teachers in the Arts and Music (perhaps applause meters… but only if we measure applause gain from concert to concert, rather than applause level?). What about vocational education?
  2. Just dump all of those teachers and all of that frivolous stuff kids don’t really need and assign each group of kids a 12 year sequence of reading and math teachers. Some have actually argued that this really should be done, especially in higher poverty and/or underperforming schools. Why, for example, should a school with inadequate math and reading scores offer instrumental music or advanced math or journalism courses? (Put down that saxophone and pick up that basic math book Mr. Parker!) The reality is that high poverty and underperforming schools in New Jersey and elsewhere already have concentrated their teaching staff on core activities to the extent that kids in poor urban schools have much less access to arts and athletics.

I personally have significant concerns over the idea that poor urban kids should have access to a string of remedial reading and math teachers over time and nothing else, but kids in affluent neighboring suburbs should be the ones with additional access to foreign languages, tennis and lacrosse teams and elite jazz ensembles (this one really irks me) and orchestras. Quite honestly, successful participation in these activities is highly relevant to college admission – at least at the competitive schools. Certainly, the affluent communities are not going to go along with dumping all of these things.

So, if we can’t test everything every year and if it is offensive to argue for dumping all areas that aren’t or can’t reasonably be evaluated, then we have a significant gap in the usefulness of VA teacher assessment.

I did this tally very quickly using 2007-08 NJ staffing files. Feel free to tally and re-tally and post alternative counts below. Note that most of the special education teachers are missing from the tally below because I’ve not yet fully recoded them for 2008. While I have done so for earlier years, those years of the staffing files don’t break out content area for MS teachers or grade level for elem teachers. About 14% of teachers in 2005 or 2006 data were special education. At a maximum, I get to about 20% of teachers as ungraded elementary and about another 5% or so potentially relevant in 2005 and 2006 for VA assessment (without ability to remove untested grades).

Main Assignment Number of Teachers % of Teachers Potentially Reliable VA Assessment No Assessment at All
Art 3,106 2.84 X
Basic Skills 1,779 1.63 X
Bilingual 697 0.64 X
Computer 917 0.84 X
Coord/Director 1,263 1.15 X
Counselors 29 0.03 X
Elem English 522 0.48
Elem Math 535 0.49
Elem Science 381 0.35
Ungraded Elem 11,308 10.33 ?
ESL 1,700 1.55 X
FCS 837 0.76 X
Grades 1 to 3 12,006 10.97
Grades 4 to 6 7,012 6.41 X
Grades 6 to 8 1,305 1.19 ?
HS English 13 0.01
HS English 5,041 4.61
HS Math 4,727 4.32
HS Science 4,391 4.01
HS Soc Studies 3,968 3.63 X
HS World Language 4,460 4.08 X
Indus Arts 1,217 1.11 X
Kindergarten 321 0.29 X
Kindergarten 3,565 3.26 X
MS Lang Arts 2,844 2.6 X
MS Math 2,439 2.23 X
MS Science 1,669 1.53 ?
MS Soc Studies 1,629 1.49 ?
MS World Language 440 0.4 X
Music 3,665 3.35 X
PE 6,963 6.36 X
Perf Arts 222 0.2 X
Preschool 1,052 0.96 X
Preschool 557 0.51 X
Principal 2,172 1.98 ?
Psychologist 1,545 1.41 X
SC Spec Educ 163 0.15 X
SC Spec Educ 6,747 6.17 X
SE RR/Inclusion 963 0.88 X
Supervisor 2,360 2.16 X
Vice Principal 1,828 1.67 X
Voc Ed 1,067 0.98 X
Total 109,433 (of about 142,000 recoded) 11.24 47.01

Okay – So New Jersey is just probably a wacky inefficient example that has way too many of those extra teachers in trivial and wasteful assignments. Well, here’s the breakout of Illinois teachers for 2008.


I could go on, and do this for Missouri, Minnesota, Wisconsin, Iowa, Washington and many others showing generally the same pattern. I chose New Jersey above  because the most recent years of NJ data actually break out the grade level assignment of most elementary teachers so we can see how many grades 1 through 3 teachers would fall outside the evaluation system.

My point here is not to try to trash VA evaluation of teachers, but rather to point out just how little – even in a practical sense – the pundits who are pitching immediate action on using VA for hiring and firing teachers and providing incentive pay have bothered to think about even the most basic issues. Not the technical and statistical issues, but really simple stuff like just how many teachers would even be evaluated under such a system. And more importantly, since this is supposedly about “incentives” – just what kind of incentives this selective evaluation might create.

Title I Does NOT make “Rich” states “Richer!”

This is one fly I keep forgetting to swat, but one that has been repeatedly advanced by the Center for American Progress with excessively crude analyses. See: http://www.americanprogress.org/issues/2009/08/title1_map.html WOW! Just look at it. Those darn rich states like Connecticut, New York and New Jersey are running away with federal funding that should be targeted to poor states like Arkansas, Alabama and Mississippi.

Two glaring omissions in this analysis undermine entirely its conclusions. First, there is the issue of regional variation in true poverty, where – because poverty thresholds used in the CAP analysis are not regionally sensitive to income variation or costs – poverty rates tend to be overstated in lower income lower cost regions. The U.S. Census Bureau has been engaged in research on this topic and released a new report last summer:

Second, the value of the Title I dollar varies significantly by location, largely as a function of the competitive wages for staff and other resources that might be purchased with those Title I dollars.

So then, how does all of this academic, trivial griping affect the CAP analysis? First, here’s a slide of the 2006-07 title I allocations per poverty pupil – same measure as CAP – by state poverty rate.

What we see here is that the small state minimum allotment does generate distorted higher amounts of T1 funding per poor child in states like North Dakota, Wyoming and Vermont.  We would also be led to believe that states like Louisiana, Mississippi, Arkansas and Tennessee are significantly disadvantaged by the formula (receiving well less than $2,000 per poor child each) and New York, Connecticut and New Jersey (hidden in the mass of points) receive around $2,000 or more (NY much more) per poor child. An abomination I say! (or at least CAP would argue).

What happens when we correct for the mis-specification of poverty, using an average of the three alternatives from the August 2009 Census Bureau paper? Well, we get:

Hmmm… Now it would appear that states like Louisiana are actually getting much more funding than New York per corrected poverty child. And Tennessee more than New Jersey! Wait – are you telling me that Title I doesn’t make these rich states richer? Yep – and I’m not even done yet.

Let’s go the next step and correct these Title I allocations per actual poor child for the regional value (based on competitive wage variation) of the Title I allocation.  Now we get:

Now we see that state’s like New Jersey, New York and especially California are actually significantly more disadvantaged by the Title I formula than states like Mississippi or Louisiana.

Look, the Title I formula certainly doesn’t produce the most logical allocations, or most equitable ones. One might also argue that it doesn’t maximize incentives for states to clean up their own act on equity or effort.

That said, there exists little excuse for excessively crude analyses which lead to such absurdly bold – AND FLAT OUT WRONG – conclusions like the conclusion that Title I makes rich states richer. Yeah – this kind of claim sounds good – makes good political rhetoric – good stump speech stuff for the absurdities of government behavior. But in this case, the CAP critique is simply wrong!

Here is a previous presentation I made on this topic before the Census working paper was available:

Baker.AERA.Title1

Let me clarify that the same issue of mis-measurement of poverty plagues urban-rural comparisons within states. Rural poverty is, in relative terms, overstated compared to urban poverty. So too are rural costs (competitive wages) lower than urban costs. So, just as it is true that Title I does not necessarily overfund “rich” states, Title I also does not necessarily overfund urban districts at the expense of rural ones. Unfortunately, I do not yet have available a finer grained adjusted poverty measure which will allow me to easily display the urban/rural issue.

Checking the Tab

As follow up to yesterday’s post on the completely fabricated and back-of-the-napkin numbers presented in The Tab,  here’s a quick simulated allocation of the $11,000 foundation + $3,000 poverty weight (applied to free or reduced lunch) + $400 per ELL/LEP child.

The Tab pretty much conceals any real changes or patterns of changes by lumping them into a summary table by groups of districts without any documentation as to how the summary stats were estimated (page 27). Above is what the district by district changes would look like. Looks pretty much like a back-of-the-napkin attempt at roughly break-even analysis. Remember, this is a proposal for the future compared against actual spending from 2007-08 – two years back now!

Specifically, the proposal would appear to reduce funding in Hartford and New Haven by greater amounts than it would increase funding in districts like New Britain and Waterbury and only similarly to the increase for Bridgeport. That is, it levels down high poverty districts as much as it levels some up – a fact concealed by the claims of a net increase of $620 per pupil in the short term. Mind you, The Tab certainly provides no evidence that districts like Hartford and New Haven are massively over-funded, as their own policy solutions would imply. Oh wait… The Tab really doesn’t rely on evidence at all. Silly me.

Just checkin the numbers – the made up numbers.

Why is it OK for Think Tanks to just make stuff up?

Something that has perplexed me for some time in my field of school finance, is why it seems to be okay for policy advocates and “Think Tanks” to just make stuff up. For example, to just make up what level of funding would be appropriate for accomplishing any particular set of goals? or to just make up a figure for how much more a child with specific educational needs requires under state school finance policy. Just “making stuff up” seems particularly problematic for “Think Tanks,” which as far as I can tell should be producing information backed by at least some degree of … Thinking? Perhaps based on some of the more reasonable thinking of the field?

This topic comes to mind today because ConnCan has just released a report (http://www.conncan.org/matriarch/documents/TheTab.pdf)    on how to fix Connecticut school funding which provides classic examples of just makin’ stuff up (page 25). The report begins with a few random charts and graphs showing the differences in funding between wealthy and poor Connecticut school districts and their state and local shares of funding. These analyses, while reasonably descriptive are relatively meaningless because they are not anchored to any well conceived or articulated explanation of “what should be.” Such a conception might be located here or even here (Chapters 13, 14 & 15 are particularly on target)!

The height of making stuff up in the report is the recommended policy solution to the problem which is never clearly articulated. There are problems in CT, but The Tab, certainly doesn’t identify them!

The supposed ideal policy solution involves a pupil-based funding formula where each pupil should receive at least $11,000 per pupil (made up), and each child in poverty (no definition provided – just a few random ideas in a footnote) should receive an additional $3,000 per pupil (also made up) and each child with limited English language proficiency should receive an additional $400 per pupil (yep… totally made up). There is minimal attempt in the report (http://www.conncan.org/matriarch/documents/TheTab.pdf) to explain why these figures are reasonable. They’re simply made up.

The authors do provide some back-of-the-napkin explanations for the numbers they made up – based on those numbers being larger than the amounts typically allocated (not necessarily true). They write off the possibility that better numbers might be derived by way of a general footnote reference to a chapter in the Handbook of Research on Education Finance and Policy by Bill Duncombe and John Yinger which actually explains methods for deriving such estimates.

The authors of The Tab conclude: “Combined with federal funding that flows on the basis of poverty and (in some cases) the English Language Learner weight of an additional $400, the $3,000 poverty weight would enable districts and schools to devote considerable resources to meeting the needs of disadvantaged students.” I’m glad they are so confident in their “made up” numbers! I, however, am less so!

It would be one thing if there was no conceptual or methodological basis for figuring out which children require more resources or how much more they might actually need. Then, I guess, you might have to make stuff up. Even then, it might be reasonable to make at least some thoughtful attempt to explain why you made up the numbers you… well… made up. But alas, such thinking seems beyond the grasp of at least some “think tanks.” Guess what? There actually are some pretty good articles out there which attempt to distill additional costs associated with specific poverty measures… like this one, by Bill Duncombe and John Yinger:

How much more does a disadvantaged student cost?

It’s not like the title of this article somehow conceals its contents, does it? Nor is the journal in which it was published (Economics of Education Review) somehow tangential to the point at hand. This paper, prepared for the National Research Council provides some additional insights into additional costs associated with poverty and methods for estimating those costs.

Rather than even attempt to argue that these figures are somehow founded in something, the authors of The Tab seem to push the point that it really doesn’t matter what these numbers are as long as the state allocates pupil-based funding.  That’s the fix! That’s what matters… not how much funding or whether the right kids get the right amounts. In fact, the reverse is true. The potential effectiveness, equity and adequacy of any decentralized weighted funding system is highly contingent upon driving appropriate levels of funding and funding differentials across schools and districts!

I’ve critiqued the notion of pupil-based funding as a panacea, here:

Review of Fund the Child: Bringing Equity, Autonomy and Portability to Ohio School Finance

Review of Shortchanging Disadvantaged Students: An Analysis of Intra-district Spending Patterns in Ohio

Review of Weighted Student Formula Yearbook 2009

Oh, and also here: http://epaa.asu.edu/epaa/v17n3/

Among other things, in each of these critiques of think-tank reports I question why it seems okay to just make up “weights” and cost figures when applying distribution formulas – either for within or between district distribution.

Just thinking… but not making stuff up!

Playing with Charter Numbers in NJ

About a week ago, I commented that charter school average performance was not much, if any different from the average performance of the poorest urban public schools. This is admittedly an oversimplified comparison, but not one I would have made had I believed it to be deceptive, which it is not – given the available data on New Jersey schools.

Here, I will walk through a more complicated though still imperfect analysis of elementary school performance in host districts and in charter schools based on data from 2004 to 2006 (data I had already compiled for related work). First, let’s begin with some descriptive characteristics of the charter schools and schools of similar grade level (elementary in this case) in their host districts based largely on school reports data from those years.

The table below shows that the data set includes 28 charter schools per year and 173 host district schools of same grade level.  The charters serve about 1,000 tested students and the host district schools about 11,000 tested students.  While the free/reduced lunch share is roughly the same between the two, the free lunch share is higher in the host district schools (these are the poorer students). These differences vary by host district and charters. Newark Charters, for example, are on average (though not all) relatively high poverty.

Note that the average free lunch share in DFG A schools, used in my previous comparison, is 63% (much higher than charters or their hosts on average).

Also higher in the host district schools are the share of children who are LEP/ELL and who are classified as having disabilities. But, the host district schools do have higher total certified salaries per pupil (compiled from state database on personnel salaries).

Slide1

Three year average scale scores are also listed, for the 2004 to 2006 period.

But, the big question is what happens when you throw this all into the mix of a statistical model to evaluate whether charters outperform host district schools, controlling for the fact that they have less needy populations, but fewer resources to work with? Again, this is a simple school level model, which does not account for individual children’s relative gains in charters (treatment effect) compared to otherwise similar children not in charters but in host district schools. It would be wonderful to be able to conduct such analyses in NJ.

This school level model includes a dummy variable for each district that is a host district, such that charter performance in the model is measured against performance of the host district of that charter. The model includes only host districts and their respective charters. The overall charter effect is essentially the average of differences between charters and hosts, across hosts (and their respective charters).

What we see in this model is that charters, on average, are no different from their hosts on the combined math and language scale scores for NJASK from 2004 to 2006.  While the statewide model of the same data shows a strong effect of cumulative salaries per pupil on outcomes, the model within host districts of charters does not – an interesting point to explore. But, other factors play out quite logically – with each student need factor statistically significantly depressing scale scores.

Slide3

So, what does this more complicated, but still not complicated enough analysis tell us? It tells us that average charter school performance from 2004 to 2006 on elementary assessments is  no different from that of average performance in other poor urban schools – specifically the host districts of those charters. It just says this in a more complicated way. Sometimes simple averages – when not deceptive – can be sufficient.

One factor that could turn the findings in favor of charters (as treatment effect) would be if the average starting performance level of charter students, compared to otherwise similar host school students, is lower than that of host school students – which could occur if there is a tendency for parents to look to charters when their children are under-performing. This appears to be the case in the Missouri data in the CREDO study noted below. But, this is unlikely to create a substantial effect.

Again, this is just playing with the numbers, albeit a more rigorous play than my previous posts – leading to the same conclusions.

For more thorough discussions of charter school research, see:

http://epicpolicy.org/think-tank/reviews

Check out specifically, the original NYC Hoxby study, and critique of it, and the CREDO 16 state study and RAND 8 state study.  Exercise caution in linking any specific findings to the New Jersey context.

Illinois Salary Gaps – Do they matter?

I picked up this article on Twitter yesterday, which seemed at first to make a veiled version of the classic “money doesn’t matter” argument, or at least that’s how the tweets and headlines were spun. The article is somewhat more thoughtful, discussing many reasons why teacher salaries vary and how those variations are largely tied to differences in taxable property wealth across Illinois school districts, but the article misses a real opportunity to shed light on some striking disparities across the state and across districts within the Chicago metro area.

http://www.chicagotribune.com/news/education/chi-teacher-salary-09-nov09,0,1639857.story

So, how might we better understand salary variation across Illinois school districts and children, and whether that variation is problematic or not? First, we know that teachers matter! Second, we know from work by Hanushek and Rivkin that the uneven distribution of teaching quality by racial composition of students can explain substantial portions of the growth in achievement gap – black-white gap – between 3rd and 8th grade. http://faculty.smu.edu/millimet/classes/eco7321/papers/hanushek%20rivkin%2002.pdf To quote:

Unequal distributions of inexperienced teachers and of racial concentrations in schools can explain all of the increased achievement gap between grades 3 and 8. (p. 1)

Further, we know from these same authors in an earlier study that:

Table 7 suggests that a school with 10 percent more black students would require about 10 percent higher salaries in order to neutralize the increased probability of leaving. (p. 38 of PDF, not numbered)

https://www.utdallas.edu/research/tsp-erc/pdf/jrnl_hanushek_2004_public_schools_lose.pdf.pdf

Recap – Two major factors determining how well kids do in school are the characteristics of the other kids in the same class and the quality of their teacher. Unfortunately, the characteristics of kids in a given class affects who typically ends up teaching that class. Classrooms with greater shares of minority children end up with less well educated, less experienced teachers. This, in combination with peer effects, produces substantial disparities in outcomes which grow over time. Salary differentials might help to offset these disparities.

As such, it would likely be quite problematic if the teacher salaries in Illinois – WITHIN ANY GIVEN LABOR MARKET – were systematically lower in districts with higher concentrations of black and/or black and hispanic children. By cursory analysis it is rather difficult to disentangle the adverse affect of salary alone. That is, you can’t just take average salaries and try to relate them to average test scores, and then conclude that salaries don’t matter, but student characteristics do. The reality is that the two simultaneously matter and interact in important ways.

In many states, salaries and overall funding are actually comparable between districts with higher minority concentrations and other districts and in some states salaries and overall funding are actually higher (though not necessarily enough higher) in higher minority concentration districts (see: http://eric.ed.gov/ERICWebPortal/custom/portlets/recordDetails/detailmini.jsp?_nfpb=true&_&ERICExtSearch_SearchValue_0=EJ718694&ERICExtSearch_SearchType_0=no&accno=EJ718694.)

OH, BUT NOT IN ILLINOIS!

In my most recent analysis of individual teacher salary data for 2004 to 2008 in Illinois, I find that for a full time teacher, at constant contractual months, same degree level and same experience, and compared to districts in the same labor market, the teacher in a school that is majority black and Hispanic children, is paid about $2,000 less per year. Further, a teacher with a masters degree makes about $8,500 more per year. And, teachers in majority minority schools in Illinois are only about 60% to 70% as likely to hold a masters degree as teachers in predominantly white schools in the same labor market in Illinois.

We also know that the dropout rate is over 7% higher in majority minority districts compared with other districts in the same labor market. The mean ACT is over 4 points lower and the mean proficiency rates on state assessments are about 20% lower in majority minority districts compared to predominantly white districts in the same labor market.

These are striking disparities. And in Illinois, unlike many other states, state policymakers have applied no financial leverage to attempt to resolve these disparities.  No harm no foul? doubtful!

Leaders and Laggards Lags!

A quick note on Center for American Progress Leaders and Laggards report.

On pages 23 & 24, this report attempts to grade state school funding systems and their level of “innovation.” But, the report pays no attention to a) whether these states actually perform well on any measures of outcomes,  b) whether these states actually fund their schools well overall, or c) whether these states actually target any of that funding to where it’s needed most.

Quite simply, this report is complete garbage – at least the finance section! One cannot possibly rate “innovation” of a state school funding system without any regard for whether that system is sufficiently and equitably funded. You can’t stimulate innovation without an investment in Research and Development or the product itself! It really is that simple.

The best Finance grades in the report are given to such education funding laggards as:

Yet, high performing states that actually fund their systems well and target resources where needed most get lousy grades (Massachusetts & New Jersey).  This  stuff is just plain silly!

====

To lighten the mood a bit, here’s Willy Wonka summarizing the Arizona school finance formula: http://www.youtube.com/watch?v=M5QGkOGZubQ

What do NJ Charter Schools Really Spend?

Getting back to the original point of my blog, this post is simply about introducing to the public discourse some actual data on NJ charter school spending. Back when I wrote my textbook on school finance, I found that DC charter schools were having to rely on private contributions to the tune of 14% of their annual operating expenses. One can obtain such information from IRS non-profit tax filings (IRS 990). I did a quick run of New Jersey Charter School IRS 990 filings for 2008, reflecting revenues and expenditures for 2007. I simply combined their tax filing information with their total expenditure information – which does include expenses for facilities.

What is most striking but not surprising is the degree of disparity among charter schools, driven substantially by differences in private fund raising.  Also important to note is that many of these schools spend well over the assumed $11,000 to $12,000 per pupil constantly spun by the media these days. I’ve not yet aligned the performance data with these new financial data, as I need to return to my actual research agenda (this particular analysis is  a part of ongoing research).

Remember also that these schools presently serve few or no special education children, making $16,000 per pupil worth well over $18,000 (assuming 15% special ed students typically at double average cost).

You might say, hey, if the public only has to subsidize $11k to $12k and private contributors pick up the rest, it’s still a bargain for taxpayers, right? Perhaps – but the necessity to rely on $3k to $5k of private contributions for each charter child educated then seriously limits the potential expansion of charter schools.

NJ Charter IRS 990Note from my previous posts and work on private schools, I have also shown that private independent day schools spend well above the average public expenditure. New Jersey private independent schools spent in 2007, an average of $25k to $30k per pupil (day schools only) with some exceeding $30k, also based on IRS 990 data.

My previous research on staffing in charter schools (based on undergraduate college selectivity of teachers) has shown that charters in some states attempt to staff their schools in ways similar to elite private academies – the private independent schools.

There is at least anecdotal evidence that some New Jersey Charter schools wish also to emulate elite private schools. For example, Ethical Community Charter School is founded by individuals previously associated with the Ethical Culture Schools of New York City, including the Fieldston School, a school where I taught for 5 years. An absolutely amazing school, which, by the way, spends well over $30,000 per child per year (even tuition is higher than that). I would argue that it will be quite difficult to emulate the ECFS schools of NYC on a mere $11k to $12k and that substantial private fundraising will be required. But private fundraising shouldn’t be required.

Good schools cost money! Sometimes a lot of money. Good education is expensive, which is not to say that all expensive education is good. My point here is that we are not going to solve our “urban education” problems on the cheap ($11k to $12k per kid), or necessarily any cheaper than what we’re spending currently. Any attempt to do so is likely to cause more harm than good.

[for those hanging on to anecdotal information about private religious school tuition as their basis for assuming good schooling can be done dirt cheap – about $3,500 per kid- please read http://www.epicpolicy.org/files/PB-Baker-PvtFinance.pdf]

The Real NJ Graduation Scam?

Bob Bowdon, of Cartel fame and E-3 make the claim that New Jersey’s poor urban districts are scamming the public and taxpayers by having overstated graduation rates. About half of poor district kids pass the HSPA test, but 85% graduate. Their brilliant solution to this problem, as I’ve noted previously, is to give kids the choice to attend charters – on the argument that charters are less likely to do such scamming?  So, here are some fun numbers.

First, the percent proficient or higher on HSPA MATH Assessments by district factor group for 2008:

Slide1

So, what we have here is that Charters (DFG R) actually had the lowest rate of kids proficient or higher on HSPA (matching my graph on previous posts, but lower here because only math is included). Yep, even lower than the poorest urban publics (DFG A). Yes, this is an average – among general ed test-takers – and averages conceal the highs… but they similarly conceal the lows.

Now, here are graduation rates for the schools by DFG:

Slide2

Wait one second. How can charters have a 97% graduation rate if only about half of the kids pass HSPA? Where’s the scam here? I thought you said that the differential between HSPA proficiency and graduation rates was supposed to be indicative of a scam? And that charters were the solution to the scam? But where is that differential bigger? Charters are lower on HSPA proficiency by a few points and are 12% higher on graduation rate? Now I’m really confused.

Okay – I’m not trying to pick on charter schools here. You guys are mostly working your butts off for a great cause, and quite honestly I don’t hear these completely absurd arguments coming from the charter leaders and teachers themselves. But the supposed “advocacy” out there on your behalf is deeply problematic. Quite honestly, if someone was out there advertising so poorly for my cause, I’d be a little concerned… or perhaps outraged.

Note to Non-Jersey readers about my casual use of Jersey terminology – DFG. In New Jersey, district factor groups or DFGs are a classification scheme that has been used for decades to characterize socio-economic features of public school districts. DFG A districts are generally poor urban districts, but many NJ poor urban districts are relatively small in total enrollment (a cluster of poor urban neighborhoods segregated from their more affluent neighbors). DFG I and J districts are affluent suburban districts. Charters are labeled “R.”

Teacher Evaluation with Value Added Measures

This month, the special issue of the journal Education Finance and Policy on value-added measurement of student outcomes was published. The table of contents is here:

http://www.mitpressjournals.org/toc/edfp/4/4

This is good stuff, authored by leading educational measurement and statistics researchers and economists. These articles provide some important cautionary tales regarding the application of value-added measures of student outcomes for teacher evaluation. Here is a policy brief with a more user friendly summary of some of the content of the special issue:

http://www.wcer.wisc.edu/publications/highlights/v19n3.pdf

Here’s a recent working paper by Jesse Rothstein, Princeton economist who also has an article in the special issue:

http://gsppi.berkeley.edu/faculty/jrothstein/published/rothstein_vam2.pdf

Here’s the concluding sentence of the abstract Rothstein’s paper:

Results indicate that even the best feasible value added models may be substantially biased, with the magnitude of the bias depending on the amount of information available for use in classroom assignments.

On average, the articles in the special issue do show some promise for using value-added assessment in teacher evaluation, with a number of really important caveats and technical stipulations.

Yes, we need access to more student assessment data with linkages to specific teachers – including the range of teachers across which middle and secondary students interact (it’s not as simple as linking the single teacher to a group of children). We need access to such data across multiple states and their assessment systems. Scaling properties of data and test noise play a major role in the precision with which one can isolate teacher or classroom level effects. We have little or no idea, for example, of the extent to which analyses using North Carolina or Texas assessment data relate to New Jersey assessment data, the statistical properties of those data and their usefulness or lack thereof for estimating teacher or classroom effects (unless there are technical papers out there on NJ tests of which I am unaware).

So, these are the main reasons we need to tear down firewalls – to advance the art, science and statistics of value added modeling, school and teacher evaluation and to uncover potential shortcomings where they exist.

Policymakers and pundits diving in head first on these issues need, quite simply, to chill out, perhaps read the special issue above and heed the advice earlier this year from the National Academy of Sciences and figure out how to do this right if we’re going to do it at all.

Diving in too quickly and doing it wrong will make it that much harder to do it right in the long run and will provide that much more ammunition for resistance.