The Principal’s Dilemma as Mock Trial: Ed Law Colleagues Please Provide Your Opinions!

The following is a hypothetical case I am using as the culminating activity in Public School Law this semester.

The Dismissal of Principal X

Principal X is principal in a local public middle school in a state that has recently adopted through legislation, articulated with greater precision in state department of education regulations, a new teacher evaluation scheme. The teacher evaluation laws and regulations now require that:

  • Any teacher who receives two sequential evaluations less than “satisfactory” shall have his/her tenure status revoked;
  • Teacher evaluations shall consist of 40 to 50% measures of student growth, where the majority shall be based on state provided metrics.
  • By regulatory decree of the State Commissioner of Education, any other measures selected by local district officials for inclusion in evaluations must be proven correlated with state approved and provided measures of student achievement growth.

Further, the state now conditions receipt of “any and all increases to state aid for local public school districts” on full compliance with statutes and regulations pertaining to teacher evaluation.

On September 20th of 2013, Principal X was provided with growth percentile data on her teachers from the prior year. Of the approximately 40 certified staff in her school, 8 received growth percentile data, two of whom achieved unsatisfactory growth percentile estimates for their students, one of whom received a second unsatisfactory rating in a row ‐Teacher Y.

In keeping with the requirement that any and all other measures used in the state approved teacher evaluations be correlated with the growth percentile measures, the principal was compelled to assign this teacher a second unsatisfactory rating, and thus compelled to revoke the tenure status of Teacher Y. She was a 10 year veteran teacher perceived by the principal and many others in the school to be one of the school’s most valuable human resources. In fact, over the past several years, the principal had relied on this teacher to take the difficult students including playing a more significant role than others in inclusion of children with disabilities in her classroom – and the teacher not only willingly, but eagerly complied.

Frustrated with the outcome of the new state teacher evaluation laws, Principal X took her case to the public and to state officials simultaneously. Without specific reference to the case in question – but via stylized example – the principal used the case of Teacher Y to illustrate how strict requirements of job action based largely on limited and problematic measures could lead to damaging decisions – decisions

1| Page

that she argued were neither in the best interest of the teachers nor the children they served, and decisions likely to negatively affect the quality of education statewide.

The principal made the case for returning discretion on issues of teacher evaluation and human resource management to local officials, including school principals. The principal’s letter led to a sympathetic uprising from community members and parents, who were quick to catch on as to which teacher was actually the basis of the principal’s hypothetical. Parents of that teacher’s students were outraged, and expressed their outrage at local board of education meetings. During this time, the local board of education maintained quiet support of the principal.

The principal had also begun to engage other principals statewide establishing a network of principals publicly proclaiming their opposition to newly adopted state teacher evaluation statutes and regulations. A web site was created, a non‐profit organization (political action organization) was formed, and the original letter of opposition to new state policies posted on the site, along with a petition for other school principals to show support for the group’s cause and/or become an official member.

State officials were less supportive and unamused by this principal’s apparent disrespect for their authority, and her “willful disobedience of existing statutes and regulations” expressed by Principal X’s stalling on submitting relevant evaluation information necessary for revocation of Teacher Y’s tenure status. Further, state officials were less than thrilled with the mounting insurrection initiated by the publicly posted letter to state officials outlining problems with the state teacher evaluation laws.

State officials released a letter to the local board of education indicating that their state aid would be frozen for the coming school year if, in fact, their rogue principal continued to stall and refuse compliance with the teacher evaluation laws. Under pressure from the board, Principal X agreed to initiate procedures that would lead to tenure revocation for Teacher Y. Instead of waiting out this process, Teacher Y chose to resign and pursue employment elsewhere.

But with mounting pressure on the local board of education from state department officials to control the growing movement among principals statewide against the teacher evaluation laws, a movement initiated by one of their most respected principals (who had received only glowing evaluations in prior years), the district board chose to dismiss Principal X, citing that the principal’s activities had distracted her from doing the job required, substantively compromised her effectiveness as a principal and significantly interfered with the ability of district officials to efficiently and effectively carry on district operations (including the uncertainty created over the district’s future state aid receipts).

Principal X is now suing the district for wrongful dismissal, arguing that the district’s dismissal is in violation of her first amendment right to express herself to the public on issues of public interest, for which she, as an informed public school employee has relevant information.

2| Page

Required Reading

Key Cases

Pickering v. Board of Education: http://www.oyez.org/cases/1960‐1969/1967/1967_510

Connick v. Myers: http://www.oyez.org/cases/1980‐1989/1982/1982_81_1251

Garcetti v. Ceballos: http://www.oyez.org/cases/2000‐2009/2005/2005_04_473

Blogs

EdJurist: Garcetti & Schools http://www.edjurist.com/garcetti‐and‐schools

EdJurist: Academic Blogging & Garcetti: http://www.edjurist.com/blog/2008/5/9/academic‐freedom­garcetti‐blogging.html

Law Reviews

Oluwole, J. O. (2007). On the Road to Garcetti: Unpick’erring Pickering and Its Progeny. Cap. UL Rev., 36, 967.

3| Page

The Perils of Economic Thinking about Human Behavior

Behavioral economics is an interesting and potentially useful field of academic inquiry. At its best, real behavioral economics attempts to address some of the concerns I raise here. But many if not most assumptions about human behavior and response to incentives are not representative of behavioral economics at its best.

Specifically,  I’m increasingly concerned with what I see as the simple-minded projection of economic thinking onto everyone and anyone else, leading to ridiculous policy recommendations – that amazingly – get taken seriously – at least by the media and punditocracy.

See, for example, Roland Fryer’s experiment on loss aversion as a strategy for incenting teachers to make sure that their students gain a few extra test score points in limited content areas. Indeed, if we pay you up front, and threaten to take your salary away if you don’t get those test score points out those kids, the data suggest  a greater likelihood of squeezing the kids for a few more points. Whether that tells us anything about the motivation and morals of teachers, or of the economists framing this argument is an entirely different question. This tells us little or nothing of the appropriate policy response. Thankfully, the policy implications of this paper were sufficiently absurd that they gained little traction.

Let’s assume classic economic assumptions about human behavior really hold steadfast and can be grossly simplified to an anything for an extra buck, or not to lose one, position. I would argue that it is perhaps economists themselves that are most stereotypical in this regard –  at least as represented in the thinking the project onto others.  In fact, I would argue that many, born out of a culture that self-selects into economic professions, are simply going out of their way to project their own thinking on others.

Further, many of these economists operate in a world where they can influence/control public policy and they too have an incentive in how they behave in this system. They are not impartial observers by any stretch of the imagination. Their goal is to use their economic research to shape public policy to their own advantage.

Put simply, just because the average morally bankrupt economist might do pretty much anything for an extra buck (or a billion), doesn’t mean the average teacher, doctor, nurse, fireman or police officer would!

This issue has been on my mind for some time, but recently came to a head when I read this completely ridiculous Washington Post article on health care policy – specifically – how to remove the incentive for hospitals and physicians resulting from surgical complications.

I should note, I come from a medical family, so some of my arguments herein are drawn from dinner table conversations (across generations), coupled with my tendency to read health policy research out of personal interest in exploring connections with education policy.

It was implied in the WaPo article… well… actually it was explicitly stated in the article that hospitals and physicians have a big financial incentive for their patients to have serious complications, leading to extended hospital stays and additional procedures.

Now, the average economist might be so morally bankrupt such that if he/she were in an operating room (OR) considering the implications of complications relative to potential earnings that they might intentionally introduce infection or other complication, but thankfully the average economist is not in the OR. Thankfully, they self-selected into economics and not medicine (likely foreseeing greater opportunity to earn more for much less work and upfront investment).

The WaPo article does make the following statement, to head off this argument:

The study does not imply that hospitals intentionally complicate surgeries to bring in more revenue.

But, I would argue that this is actually a rather half-hearted disclaimer (to a half-assed argument) for an article that very much implies just that.

Certainly the economists’ policy response – how to employ crude economic assumptions of human behavior to fix this dreadful perverse incentive – implies that cutting off this financial benefit for malpractice would improve hospital and physician behavior [meanwhile conflating the hospital and physician incentives & roles in the various related processes]. Here is the policy solution recommended by the economists cited in the WaPo article:

If hospitals receive a set amount for every heart surgery they perform, for example, they suddenly have an incentive to reduce complications — they know the extra medical spending will come out of their own budget.

Lost in the economists’ reasoning here are a) the potential longer term financial and career implications to the physician repeatedly entangled in litigation over post-surgical complications, b) and the stress/mental toll on the physician arising from managing complications in tense moments in the OR.

Indeed this is anecdotal, but I’ve not met a physician – surgeon or anesthesiologist – who prefers a day when things go bad in the OR – or would be likely to see dollar signs in those moments of stress. What kind of sick bastard even thinks that way? Well, perhaps the average economist does.

Economists rarely – uh… NEVER face comparable professional stress to managing a patient’s life on the edge – even when they make a massively stupid spreadsheet error stimulating economic turmoil across the globe. Nor do they pay hefty malpractice premiums to shield themselves from such egregious malpractice (despite measurable financial damages). I would assert that the economist never faces the stress of having to care for a classroom of 20 to 40, 5 to 15 year old kids, whose immediate safety and well-being, as well as their long term futures is on the line.  This is in part, why they get away with such ludicrous thinking.

It’s all freakin’ game (Freakin’ used here in a technical freakonomics sense)… a game of playing with big data – several layers removed from reality – from people – from real human consequences.

Perhaps that’s the central issue… even more so than economists’ financial self-interests?

Taken in perspective, it’s a fun game and a pretty cushy lifestyle to have opportunity to ponder policy implications of big data, as long as we don’t start thinking that what we do is so freakin’ important and indispensable and as long as we understand where we sit in this big messy puzzle of human behavior and incentives.

Of course, the other interesting piece here is the leap we often see these days between what the study behind the headlines actually said, and the resulting spin in the media headlines. We also often see the economists themselves engaging in the spin. This was equally true in the famed Chetty, Rockoff, Friedman Fire Teachers First, Ask Questions Later study.

For example, here’s what the original study – in the Journal of the American Medical Association – on reimbursements associated with complications actually said:

Depending on payer mix, many hospitals have the potential for adverse near-term financial consequences for decreasing postsurgical complications.

It takes one hell of a leap of logic to get from this measured finding to the policy recommendation above.  It takes projecting economists thinking –amoral greed – onto all actors involved. It also takes ignoring entirely a multitude of contextual factors and perverse consequences (economist thinking – first, we assume none of that stuff exists). Indeed, many complications relate to preexisting conditions and/or overall health of the incoming patient. Do we really want to incent risk aversion? (avoiding those far more likely to have complications?). Well, if it leads to lower premiums for and taxes paid by economists, then perhaps?

Tangentially (or not?), there is an equally ill-conceived movement afoot to apply to healthcare management the brilliance of what we have supposedly learned from measuring teacher effectiveness with value-added models, as explained in this policy brief from Mathematica. Notably, I tend to think Mathematica does pretty good work on education policy (better than most. See here, but for more critical perspective, see here) Put in its best light, this policy brief is merely Mathematica researchers engaging in another I’ve got a Hammer… where’s the freakin’ nail exercise.

Put in the light of economic thinking about human behavior – which many economists prefer to project on all others, the incentive here is for Mathematica to broaden its market, gaining contracts to develop value-added metrics for health care systems and for state and Federal government – to ultimately be used in reducing payments for healthcare, and reducing the tax burden and healthcare premiums paid by Mathematica researchers – their funders and their peers. It’s a win/win. More contracts and higher income, and lower taxes and health benefits expenses (not costs, but expenses*).

That is, as long as they are never in need of surgery.

====

*Cost  reduction implies that quality of service remains constant, whereas expenditure reduction may lead to service quality reduction.

Revisiting the Complexities of Charter Funding Comparisons

This Education Week Post today rather uncritically summarized a recently published article based on an earlier report on charter school spending “gaps.” I’ve not had a chance to dig into this updated study yet, but the Ed Week post also referred to an earlier study from Ball State University which I have critiqued on multiple occasions. Importantly, my previous critiques of this study point to the complexities of making these comparisons appropriately.  Here is one version of my critique of the Ball State study, which appears in Footnote 22, page 49 of this study: http://nepc.colorado.edu/files/rb-charterspending_0.pdf

A study frequently cited by charter advocates, authored by researchers from Ball State University and Public Impact, compared the charter versus traditional public school funding deficits across states, rating states by the extent that they under-subsidize charter schools. The authors identify no state or city where charter schools are fully, equitably funded.

But simple direct comparisons between subsidies for charter schools and public districts can be misleading because public districts may still retain some responsibility for expenditures associated with charters that fall within their district boundaries or that serve students from their district. For example, under many state charter laws, host districts or sending districts retain responsibility for providing transportation services, subsidizing food services, or providing funding for special education services. Revenues provided to host districts to provide these services may show up on host district financial reports, and if the service is financed directly by the host district, the expenditure will also be incurred by the host, not the charter, even though the services are received by charter students.

Drawing simple direct comparisons thus can result in a compounded error: Host districts are credited with an expense on children attending charter schools, but children attending charter schools are not credited to the district enrollment. In a per-pupil spending calculation for the host districts, this may lead to inflating the numerator (district expenditures) while deflating the denominator (pupils served), thus significantly inflating the district’s per pupil spending. Concurrently, the charter expenditure is deflated.
Correct budgeting would reverse those two entries, essentially subtracting the expense from the budget calculated for the district, while adding the in-kind funding to the charter school calculation. Further, in districts like New York City, the city Department of Education incurs the expense for providing facilities to several charters. That is, the City’s budget, not the charter budgets, incur another expense that serves only charter students. The Ball State/Public Impact study errs egregiously on all fronts, assuming in each and every case that the revenue reported by charter schools versus traditional public schools provides the same range of services and provides those services exclusively for the students in that sector (district or charter).

Charter advocates often argue that charters are most disadvantaged in financial comparisons because charters must often incur from their annual operating expenses, the expenses associated with leasing facilities space. Indeed it is true that charters are not afforded the ability to levy taxes to carry public debt to finance construction of facilities. But it is incorrect to assume when comparing expenditures that for traditional public schools, facilities are already paid for and have no associated costs, while charter schools must bear the burden of leasing at market rates – essentially and “all versus nothing” comparison. First, public districts do have ongoing maintenance and operations costs of facilities as well as payments on debt incurred for capital investment, including new construction and renovation. Second, charter schools finance their facilities by a variety of mechanisms, with many in New York City operating in space provided by the city, many charters nationwide operating in space fully financed with private philanthropy, and many holding lease agreements for privately or publicly owned facilities.

New York City is not alone it its choice to provide full facilities support for some charter school operators (http://www.thenotebook.org/blog/124517/district-cant-say-how-many-millions-its-spending-renaissance-charters). Thus, the common characterization that charter schools front 100% of facilities costs from operating budgets, with no public subsidy, and traditional public school facilities are “free” of any costs, is wrong in nearly every case, and in some cases there exists no facilities cost disadvantage whatsoever for charter operators.

Baker and Ferris (2011) point out that while the Ball State/Public Impact Study claims that charter schools in New York State are severely underfunded, the New York City Independent Budget Office (IBO), in more refined analysis focusing only on New York City charters (the majority of charters in the State), points out that charter schools housed within Board of Education facilities are comparably subsidized when compared with traditional public schools (2008-09). In revised analyses, the IBO found that co-located charters (in 2009-10) actually received more than city public schools, while charters housed in private space continued to receive less (after discounting occupancy costs). That is, the funding picture around facilities is more nuanced that is often suggested.

Batdorff, M., Maloney, L., May, J., Doyle, D., & Hassel, B. (2010). Charter School Funding: Inequity Persists. Muncie, IN: Ball State University.

NYC Independent Budget Office (2010, February). Comparing the Level of Public Support: Charter Schools versus Traditional Public Schools. New York: Author, 1.

NYC Independent Budget Office (2011). Charter Schools Housed in the City’s School Buildings get More Public Funding per Student than Traditional Public Schools. New York: Author. Retrieved April 24, 2012, from http://ibo.nyc.ny.us/cgi-park/?p=272.

NYC Independent Budget Office (2011). Comparison of Funding Traditional Schools vs. Charter Schools: Supplement. New York: Author .Retrieved April 24, 2012, from http://www.ibo.nyc.ny.us/iboreports/chartersupplement.pdf.

Note: The average “capital outlay” expenditure of public school districts in 2008-09 was over $2,000 per pupil in New York State, nearly $2,000 per pupil in Texas and about $1,400 per pupil in Ohio. Based on enrollment weighted averages generated from the U.S. Census Bureau’s Fiscal Survey of Local Governments, Elementary and Secondary School Finances 2008-09 (variable tcapout): http://www2.census.gov/govs/school/elsec09t.xls

Friday AM Graphs: Just how biased are NJ’s Growth Percentile Measures (school level)?

New Jersey finally released the data set of its school level growth percentile metrics. I’ve been harping on a few points on this blog this week.

SGP data here: http://education.state.nj.us/pr/database.html

Enrollment data here: http://www.nj.gov/education/data/enr/enr12/stat_doc.htm

First, that the commissioner’s characterization that the growth percentiles necessarily fully take into account student background is a completely bogus and unfounded assertion.

Second, that it is entirely irresponsible and outright reckless that they’ve chosen not even to produce technical reports evaluating this assertion.

Third, that growth percentiles are merely individual student level descriptive metrics that simply have no place in the evaluation of teachers, since they are not designed (by their creator’s acknowledgement) for attribution of responsibility for that student growth.

Fourth, that the Gates MET studies provide absolutely no validation of New Jersey’s choice to use SGP data in the way proposed regulations mandate.

So, this morning I put together four quick graphs of the relationship between school level percent free lunch and median SGPs in language arts and math and school level 7th grade proficiency rates and median SGPs in language arts and math. Just how bad is the bias in the New Jersey SGP/MGP data?  Well, here it is! (actually, it was bad enough to shock me)

First, if you are a middle school with higher percent free lunch, you are, on average likely to have a lower growth percentile rating in Math. Notably, the math ASK assessment has significant ceiling effect leading into middle grades, perhaps weakening this relationship. (more on this at a later point)Slide1

If your are a middle school with higher percent free lunch, you are, on average, likely to have a lower growth percentile rating in English Language Arts. This relationship is actually even more biased than the math relationship (uncommon for this type of analysis), likely because the ELA assessment suffers less ceiling effect problem.

Slide2As with many if not most SGP data, the relationship is actually even worse when we look at the correlation with average performance level of the school, or peer group. If your school has higher proficiency rates to begin with, your school will quite likely have a higher growth percentile ranking:

Slide3

The same applies for English Language Arts:

Slide4

Quite honestly these the worst – most biased – school level growth data I think I’ve ever seen.

They are worse than New York State.

They are much worse than New York City.

And they are worse than Ohio.

And this is just a first cut at them. I suspect that if I have actual initial scores or even school level scale scores, the relationship between those scores and growth percentile is even stronger. But will test that when opportunity presents itself.

Further, because the bias is so strong at the school level – it is likely also quite strong at the teacher level.

New Jersey’s school level MGPs are highly unlikely to be providing any meaningful indicator of the actual effectiveness of teachers, administrators and practices of New Jersey schools.  Rather, by conscious choice to ignore contextual factors of schooling (be it the vast variations in the daily lives of individual children, or the difficult to measure power of peer group context, and various  other social contextual factors), New Jersey’s growth percentile measures fail miserably.

No school can be credibly rated as effective or not based on these data, nor can any individual teacher be cast as necessarily effective or ineffective.

And this not at all unexpected.

Additional Graphs: Racial Bias

Slide5

Slide6

Just for fun, here’s a multiple regression model which yields additional factors that are statistically associated with school level MGPs. First and foremost, these factors explain over 1/3 of the variation in Language Arts MGPs. That is, Language Arts MGPs seem heavily contingent upon a) student demographics, b) location and c) grade range of school.  In other words, if we start using these data as a basis for de-tenuring teachers, we will likely be detenuring teachers quite unevenly with respect to a) student demographics, b) location and c) grade range… despite having little evidence that we are actually validly capturing teacher effectiveness – and substantial implication here that we are, in fact, NOT.

Patterns for math aren’t much different. Less variance is explained, again, I suspect because of the strong ceiling effect on math assessments in the upper elementary/middle grades. There appears to be a charter school positive effect in this regression, but I remain too suspicious of attaching any meaningful conclusions to these data. Besides, if we assert this charter effect to be true as a function of these MGPs being somehow valid, then we’d have to accept that charters like Robert Treat in Newark are doing a particularly poor job (very low MGP either compared to similar demographic schools, or similar average performance level schools).

School Level Regression of Predictors of Variation in MGPs

school mgp regression

*p<.05, **p<.10

At this point, I think it’s reasonable to request that the NJDOE turn over masked (removing student identifiers) versions of their data… the student level SGP data (with all relevant demographic indicators), matched to teachers, attached to school IDs, and also including certifying institutions of each teacher.  These data require thorough vetting at this point as it would certainly appear that they are suspect as a school evaluation tool. Further, any bias that becomes apparent to this degree at the school level – which is merely an aggregation of teacher/classroom level data – indicates that these same problems exist in the teacher level data. Given the employment consequences here, it is imperative that NJDOE make these data available for independent review.

Until these data are fully disclosed (not just their own analyses of them, which I expect to be cooked up any day now), NJDOE and the Board of Education should immediately cease moving forward on using these data either for any consequential decisions either for schools or individual teachers. And if they do not, school administrators, local boards of education and individual teachers and teacher preparation institutions (which are also to be rated by this shoddy information) should JUST SAY NO!

A few more supplemental analyses

Slide1

Slide2

Slide3

Slide4

 

Briefly Revisiting the Central Problem with SGPs (in the creator’s own words)

When I first criticized the use of SGPs for teacher evaluation in New Jersey, the creator of the Colorado Growth Model responded with the following statement:

Unfortunately Professor Baker conflates the data (i.e. the measure) with the use. A primary purpose in the development of the Colorado Growth Model (Student Growth Percentiles/SGPs) was to distinguish the measure from the use: To separate the description of student progress (the SGP) from the attribution of responsibility for that progress.

http://www.ednewscolorado.org/voices/student-growth-percentiles-and-shoe-leather

I responded here.

Let’s parse this statement one more time. The goal, of the SGP approach, as applied in the Colorado Growth Model and subsequently in other states is to:

…separate the description of student progress (the SGP) from the attribution of responsibility for that progress.

To evaluate the effectiveness of a teacher on influencing student progress, one must certainly be able to attribute responsibility for that progress to the teacher. If SGP’s aren’t designed to attribute that responsibility, then they aren’t designed for evaluating teacher effectiveness, and thus aren’t a valid factor for determining whether a teacher should have his/her tenure revoked on the basis of their ineffectiveness.

It’s just that simple!

Employment lawyers, save the quote and link above for cross examination of Dr. Betebenner when teachers start losing their tenure status and/or are dismissed primarily on the basis of his measures – which by his own recognition – are not designed to attribute responsibility for student growth to them (the teachers) or any other home, school or classroom factor that may be affecting that growth.

(Reiterating again that while value added models do attempt to isolate teacher effect, they just don’t do a very good job at it).

 

 

 

 

On Misrepresenting (Gates) MET to Advance State Policy Agendas

In my previous  post I chastised state officials for their blatant mischaracterization of metrics to be employed in teacher evaluation. This raised (in twitter conversation) the issue of the frequent misrepresentation of findings from the Gates Foundation Measures of Effective Teaching Project (or MET). Policymakers frequently invoke the Gates MET findings as providing broad based support for however they might choose to use, whatever measures they might choose to use (such as growth percentiles).

Here is one example in a recent article from NJ Spotlight (John Mooney) regarding proposed teacher evaluation regulations in New Jersey:

New academic paper: One of the most outspoken critics has been Bruce Baker, a professor and researcher at Rutgers’ Graduate School of Education. He and two other researchers recently published a paper questioning the practice, titled “The Legal Consequences of Mandating High Stakes Decisions Based on Low Quality Information: Teacher Evaluation in the Race-to-the-Top Era.” It outlines the teacher evaluation systems being adopted nationwide and questions the use of SGP, specifically, saying the percentile measures is not designed to gauge teacher effectiveness and “thus have no place” in determining especially a teacher’s job fate.

The state’s response: The Christie administration cites its own research to back up its plans, the most favored being the recent Measures of Effective Teaching (MET) project funded by the Gates Foundation, which tracked 3,000 teachers over three years and found that student achievement measures in general are a critical component in determining a teacher’s effectiveness.

I asked colleague Morgan Polikoff of the University of Southern California for his comments. Note that Morgan and I aren’t entirely on the same page on the usefulness of even the best possible versions of teacher effect (on test score gain) measures… but we’re not that far apart either.  It’s my impression that Morgan believes that better estimated measures can be more valuable – more valuable than I perhaps think they can be in policy decision making. My perspective is presented here (and Morgan is free to provide his).  My skepticism in part arises from my perception that there is neither interest among or incentive for state policymakers to actually develop better measures (as evidenced in my previous post). And that I’m not sure some of the major issues can ever be resolved.

That aside, here are Morgan Polikoff’s comments regarding misrepresentation of the Gates MET findings – in particular, as applied to states adopting student growth percentile measures:

As a member of the Measures of Effective Teaching (MET) project research team, I was asked by Bruce to pen a response to the state’s use of MET to support its choice of student growth percentiles (SGPs) for teacher evaluations. Speaking on my behalf only (and not on behalf of the larger research team), I can say that the MET project says nothing at all about the use of SGPs. The growth measures used in the MET project were, in fact, based on value-added models (VAMs) (http://www.metproject.org/downloads/MET_Gathering_Feedback_Research_Paper.pdf). The MET project’s VAMs, unlike student growth percentiles, included an extensive list of student covariates, such as demographics, free/reduced-price lunch, English language learner, and special education status.

Extrapolating from these results and inferring that the same applies to SGPs is not an appropriate use of the available evidence. The MET results cannot speak to the differences between SGP and VAM measures, but there is both conceptual and empirical evidence that VAM measures that control for student background characteristics are more conceptually and empirically appropriate (link to your paper and to Cory Koedel’s AEFP paper). For instance, SGP models are likely to result in teachers teaching the most disadvantaged students being rated the poorest (cite Cory’s paper). This may result in all kinds of negative unintended consequences, such as teachers avoiding teaching these kinds of students.

In short, state policymakers should consider all of the available evidence on SGPs vs. VAMs, and they should not rely on MET to make arguments about measures that were not studied in that work.

Morgan

Citations:

Baker, B.D., Oluwole, J., Green, P.C. III (2013) The legal consequences of mandating high stakes decisions based on low quality information: Teacher evaluation in the race-to-the-top era. Education Policy Analysis Archives, 21(5). This article is part of EPAA/AAPE’s Special Issue On Value-Added: What America’s Policymakers Need to Know and Understand, Guest Edited by Dr. Audrey Amrein-Beardsley and Assistant Editors Dr. Clarin Collins, Dr. Sarah Polasky, and Ed Sloat. Retrieved [date], from http://epaa.asu.edu/ojs/article/view/1298

Ehlert, M., Koedel, C., Parsons, E., & Podgursky, M. (2012). Selecting Growth Measures for School and Teacher Evaluations. http://ideas.repec.org/p/umc/wpaper/1210.html

(Updated alternate version:

http://economics.missouri.edu/working-papers/2012/WP1210_koedel.pdf)

 

Who will be held responsible when state officials are factually wrong? On Statistics & Teacher Evaluation

While I fully understand that state education agencies are fast becoming propaganda machines, I’m increasingly concerned with how far this will go.  Yes, under NCLB, state education agencies concocted completely wrongheaded school classification schemes that had little or nothing to do with actual school quality, and in rare cases, used those policies to enforce substantive sanctions on schools. But, I don’t recall many state officials going to great lengths to prove the worth – argue the validity – of these systems. Yeah… there were sales-pitchy materials alongside technical manuals for state report cards, but I don’t recall such a strong push to advance completely false characterizations of the measures. Perhaps I’m wrong. But either way, this brings me to today’s post.

I am increasingly concerned with at least some state officials’ misguided rhetoric promoting policy initiatives built on information that is either knowingly suspect, or simply conceptually wrong/inappropriate.

Specifically, the rhetoric around adoption of measures of teacher effectiveness has become driven largely by soundbites that in many cases are simply factually WRONG.

As I’ve explained before…

  • With value added modeling, which does attempt to parse statistically the relationship between a student being assigned to teacher X and that students achievement growth, controlling for various characteristics of the student and the student’s peer group, there still exists a substantial possibility of random-error based mis-classification of the teacher or remaining bias in the teacher’s classification (something we didn’t catch in the model affected that teacher’s estimate). And there’s little way of knowing what’s what.
  • With student growth percentiles, there is no attempt to parse statistically the relationship between a student being assigned a particular teacher and the teacher’s supposed responsibility for that student’s change among her peers in test score percentile rank.

This article explains these issues in great detail.

And this video may also be helpful.

Matt Di     Carlo has written extensively about the question of whether and how well value-added modes actually accomplish their goal of fully controlling for student backgrounds.

Sound Bites don’t Validate Bad or Wrong Measures!

So, let’s take a look at some of the rhetoric that’s flying around out there and why and how it’s WRONG.

New Jersey has recently released its new regulations for implementing teacher evaluation policies, with heavy reliance on student growth percentile scores, ultimately aggregated to the teacher level as median growth percentiles. When challenged about whether those growth percentile scores will accurately represent teacher effectiveness, specifically for teachers serving kids from different backgrounds, NJ Commissioner Christopher Cerf explains:

“You are looking at the progress students make and that fully takes into account socio-economic status,” Cerf said. “By focusing on the starting point, it equalizes for things like special education and poverty and so on.” (emphasis added)

http://www.wnyc.org/articles/new-jersey-news/2013/mar/18/everything-you-need-know-about-students-baked-their-test-scores-new-jersy-education-officials-say/

Here’s the thing about that statement. Well, two things. First, the comparisons of individual students don’t actually explain what happens when a group of students is aggregated to their teacher and the teacher is assigned the median student’s growth score to represent his/her effectiveness, where teacher’s don’t all have an evenly distributed mix of kids who started at similar points (to other teachers). So, in one sense, this statement doesn’t even address the issue.

More importantly, however, this statement is simply WRONG!

There’s little or no research to back this up, but for early claims of William Sanders and colleagues in the 1990s in early applications of value added modeling which excluded covariates. Likely, those cases where covariates have been found to have only small effects are cases in which those effects are drowned out by noise or other bias resulting from underlying test scaling (or re-scaling) issues – or alternatively, crappy measurement of the covariates. Here’s an example of the stepwise effects of adding covariates on teacher ratings.

Consider that one year’s assessment is given in April. The school year ends in late June. The next year’s test is given the next April. First, and tangential (to the covariate issue… but still important) there are approximately two months of instruction given by the prior year’s teacher that are assigned the current year’s teacher. Beyond that, there are a multitude of things that go on outside of the few hours a day where the teacher has contact with a child, that influence any given child’s “gains” over the year, and those things that go on outside of school vary widely by children’s economic status. Further, children with certain life experiences on a continued daily/weekly/monthly basis are more likely to be clustered with each other in schools and classrooms.

With annual test scores – differences in summer experiences (slide 20) which vary by student economic background matter – differences in home settings and access to home resources matters – differences in access to outside of school tutoring and other family subsidized supports may matter and depend on family resources.  Variations in kids’ daily lives more generally matter (neighborhood violence, etc.) and many of those variations exist as a function of socio-economic status.

Variations in peer group with whom children attend school matters, and also varies by socio-economic status, neighborhood structure, conditions, and varies by socioeconomic status of not just the individual child, but the group of children. (citations and examples available in this slide set)

In short, it is patently false to suggest that using the same starting point “fully takes into account socio-economic status.”

It’s certainly false to make such a statement about aggregated group comparisons – especially while never actually conducting or producing publicly any analysis to back such a ridiculous claim.

For lack of any larger available analysis of aggregated (teacher or school level) NJ growth percentile data, I stumbled across this graph from a Newark Public Schools presentation a short while back.

NPS SGP Bias

http://www.njspotlight.com/assets/12/1212/2110

Interestingly, what this graph shows is that the average score level in schools is somewhat positively associated with the median growth percentile, even within Newark where variation is relatively limited. In other words, schools with higher average scores appear to achieve higher gains. Peer group effect? Maybe. Underlying test scaling effect? Maybe. Don’t know. Can’t know.

The graph provides another dimension that is also helpful. It identifies lower and higher need schools – where “high need” are the lowest need in the mix. They have the highest average scores, and highest growth percentiles. And this is on the English/language arts assessment, where Math assessments tend to reveal stronger such correlations.

Now, state officials might counter that this pattern actually occurs because of the distribution of teaching talent… and has nothing to do with model failure to capture differences in student backgrounds. All of the great teachers are in those lower need, higher average performing schools! Thus, fire the others, and they’ll be awesome too! There is no basis for such a claim given that the model makes no attempt beyond prior score to capture student background.

Then there’s New York State, where similar rhetoric has been pervasive in the state’s push to get local public school districts to adopt state compliant teacher evaluation provisions in contracts, and to base those evaluations largely on state provided growth percentile measures. Notably, New York State unlike New Jersey actually realized that the growth percentile data required adjustment for student characteristics. So they tried to produce adjusted measures. It just didn’t work.

In a New York Post op-ed, the Chancellor of the Board of Regents opined:

The student-growth scores provided by the state for teacher evaluations are adjusted for factors such as students who are English Language Learners, students with disabilities and students living in poverty. When used right, growth data from student assessments provide an objective measurement of student achievement and, by extension, teacher performance. http://www.nypost.com/p/news/opinion/opedcolumnists/for_nyc_students_move_on_evaluations_EZVY4h9ddpxQSGz3oBWf0M

So, what’s wrong with that? Well… mainly… that it’s… WRONG!

First, as I elaborate below, the state’s own technical report on their measures found that they were in fact not an unbiased measure of teacher or principal performance:

Despite the model conditioning on prior year test scores, schools and teachers with students who had higher prior year test scores, on average, had higher MGPs. Teachers of classes with higher percentages of economically disadvantaged students had lower MGPs. (p. 1) https://schoolfinance101.com/wp-content/uploads/2012/11/growth-model-11-12-air-technical-report.pdf

That said, the Chancellor has cleverly chosen her words. Yes, it’s adjusted… but the adjustment doesn’t work. Yes, they are an objective measure. But they are still wrong. They are a measure of student achievement. But not a very good one.

But they are not by any stretch of the imagination, by extension, a measure of teacher performance. You can call them that. You can declare them that in regulations. But they are not.

To ice this reformy cake in New York, the Commissioner of Education has declared in letters to individual school districts regarding their evaluation plans, that any other measure they choose to add along side the state growth percentiles must be acceptably correlated with the growth percentiles:

The department will be analyzing data supplied by districts, BOCES and/or schools and may order a corrective action plan if there are unacceptably low correlation results between the student growth subcomponent and any other measure of teacher and principal effectiveness… https://schoolfinance101.wordpress.com/2012/12/05/its-time-to-just-say-no-more-thoughts-on-the-ny-state-tchr-eval-system/

Because, of course, the growth percentile data are plainly and obviously a fair, balanced objective measure of teacher effectiveness.

WRONG!

But it’s better than the Status Quo!

The standard retort is that marginally flawed or not, these measures are much better than the status quo. ‘Cuz of course, we all know our schools suck. Teachers really suck. Principals enable their suckiness.  And pretty much anything we might do… must suck less.

WRONG – it is absolutely not better than the status quo to take a knowingly flawed measure, or a measure that does not even attempt to isolate teacher effectiveness, and use it to label teachers as good or bad at their jobs. It is even worse to then mandate that the measure be used to take employment action against the employee.

It’s not good for teachers AND It’s not good for kids. (noting the stupidity of the reformy argument that anything that’s bad for teachers must be good for kids, and vice versa)

On the one hand, these ridiculous rigid, ill-conceived, statistically and legally inept and morally bankrupt policies will most certainly lead to increased, not decreased litigation over teacher dismissal.

On the other hand… The anything is better than the status quo argument is getting a bit stale and was pretty ridiculous to begin with.  Jay Matthews of the Washington Post acknowledged his preference for a return toward the status quo (suggesting different improvements) in a recent blog post, explaining:

We would be better off rating teachers the old-fashioned way. Let principals do it in the normal course of watching and working with their staff. But be much more careful than we have been in the past about who gets to be principal, and provide much more training.

In closing, the ham-fisted argument of the anti-status quo argument, as applied to teacher evaluation, is easily summarized as follows:

Anything > Status Quo

Where the “greater than” symbol implies “really freakin’ better than… if not totally awesome… wicked awesome in fact,” but since it’s all relative, it would have be “wicked awesomer.”

Because student growth measures exists and purport to measure student achievement growth which is supposed to be a teacher’s primary responsibility, it therefore counts as “something,” which is a subclass of “anything” and therefore it is better than the “status quo.” That is:

Student Growth Measures = “something”

Something ⊆ Anything (something is a subset of anything)

Something > Status Quo

Student Growth Measures > Current Teacher Evaluation

Again, where “>”  means “awesomer” even though we know that current teacher evaluation is anything but awesome.

It’s just that simple!

And this is the basis for modern education policymaking?

The disturbing language and shallow logic of Ed Reform: Comments on “Relinquishment” & “Sector Agnosticism”

Two buzz phrases have been somewhat quietly floating around reformyland of late, for at least a year or so. I suspect that many have not even picked up on these buzz phrases/words.  They are somewhat inner circle concepts in reformyland. The first is the notion of the great relinquisher (a seemingly bizarre contradiction indeed… to be great at surrendering… but I believe that’s the point). The second is the idea that we all must learn to be sector agnostics. That is, we all must stand behind the provision of a system of great schools as logical replacement for existing school systems and that this system of great schools might be provided by any sector – public/government, charter, private non-profit, private for profit. After all, it doesn’t matter how we provide them, as long as they are great schools. Who can argue with that?

Linking these two conceptions, the great relinquishers – primarily public officials perceived as otherwise self-interested bureaucrats – must learn to relinquish their self-interested stronghold on publicly financed schooling to alternative providers.  Among inner circle reformers, these ideas are treated as somehow ground breaking, deep intellectual thoughts about re-envisioning schooling. But in reality, they are anything but.

On Relinquishers & Sector Agnosticism

Some abbreviated backdrop on the relinquisher notion.  I converse (constructively) on occasion via e-mail with Neerav Kingsland who promotes this particular notion. For those who don’t know Neerav, he’s a Yale Law grad who completed a Broad Residency, and is currently CEO for New Schools for New Orleans. Thus, as I interpret it, he derives his core arguments largely on his perception of the (highly debatable) successes of post-Katrina New Orleans.  That in mind, and with all due respect to Neerav, I have grave concerns about what he refers to as the movement toward “relinquishment” or creating a culture of “relinquishers” among current public officials regarding the provision of the public good of schooling (differing substantively from public schooling.)

Neerav introduced the concept of Relinquishers in a letter he wrote to urban (not all, just “urban”) superintendents in Education Week:

Before I begin in full, let me say this: Superintendents, over the years I’ve begun to believe that your identities–how each of you perceives your professional charge–are often misguided. In my experience, most of you view yourselves as system reformers–leaders who can make the current educational system much better. For the sake of the letter, let’s call you, well, Reformers. With great diligence, you fight to make our government-operated system better.

But let me suggest another identity–one whose charge is to return power, in a thoughtful manner, back to parents and educators. Let’s call these types of superintendents Relinquishers. With great diligence, these superintendents attempt to transfer power away from a centralized bureaucracy.

Both Reformers and Relinquishers possess noble aims, but only one group, I think, possesses a sound strategy.

Superintendents, in the rest of this letter I hope to convince you to become Relinquishers. Specifically, I will advocate that you return power to parents and educators through the creation of charter school districts, which are the most politically acceptable mechanisms for empowering educators. (my emphasis)

Let’s start by taking the word “relinquish” literally for a moment. A quick synonym search in Microsoft Word yields: Surrender, Abandon, Renounce, Resign

The implication here is that public officials must “surrender” or “abandon” or “renounce” their schools, handing them over largely to private managers of charter schools (note that Neerav Kingsland has suggested that charter operators are the “politically acceptable” choice, leaving for others including Smarick to consider conventional private schooling and voucher models). Yeah… I get that this is an interesting notion – to suggest that there is some nobility is declaring defeat and handing control over to those who might be able to play a positive role. I get that. But I find this use, and this framing rather disturbing.

This is not to suggest that I don’t believe that many local public school districts, large or small, need work (some, a hell of a lot of work) on how they interact with their local communities and how they balance stakeholder interests (responsiveness to parents/students, etc.). That’s an ongoing concern in any public or private sector business, with differing structural/governance issues involved in public governance. This is also not to suggest that public officials should never look to other sectors for appropriately contracted, sufficiently regulated support. But “relinquishment” is an extreme perversion of this notion, especially when we start considering relinquishment of the system as a whole – Surrendering, abandoning, renouncing any and all role for public governance and centralized public policy.

Now for this notion of “sector agnosticism” – In his book The Urban School System of the Future and in several tweets and blog posts, former deputy commissioner of Education of New Jersey, Andrew Smarick promotes the reformy religion of what he refers to as Sector Agnosticism. A brief explanation is provided in a recent education week post:

Smarick: “Second, we need to have a three-sector accountability system that treats similarly district public schools, charter public schools, and private schools; we must focus on school results, not school operator. I call this “sector agnosticism;” in other words, we shouldn’t care who runs a school as long as it is superb.”

In the 1990s, when this idea arguably first gained some momentum (summarized in Paul Hill’s book Reinventing Public Education), I was actually a pretty big fan of the idea – which consisted primarily of finding ways to employ private contractors through performance contracting to improve urban schools. Heck, my own first conference paper ever was on the issue of private management of public schools, at a time when I thought there might be great hope for such strategies. Unfortunately, the self-interest of the (publicly traded, for profit) private manager (who eventually fell into financial collapse) to extract as much revenue as possible from the urban district (Baltimore) coupled with their outright disinterest in, and obstruction of having their outcomes measured, started giving me doubts. How could they possible show an efficiency advantage (doing more with less) if they managed to game their budget allocations to their advantage and then wouldn’t provide evidence of results?

Unfortunately, I wrongly assumed things would get better as the industry evolved. Further, over time, as I completed graduate work studying education finance and policy and became reasonably well versed in school law and education governance (teaching it at the graduate level for over a decade & writing/publishing numerous co-authored articles in law review journals) I became more acutely aware of the potential pitfalls of taking an uniformed leap into sector agnosticism.

Defining superbitude?

First, let’s take Smarick’s sound bite notion that it should matter as long as the school is “superb.” Even with a narrow, test score or graduation & post-secondary matriculation-based measure of “superbitude,” neither charter nor private schools are revealing any decisive edge, holding student characteristics or access to resources constant. Rather, as one might logically expect, these less regulated sectors merely produce greater variation around largely the same mean (if comparing similar students). Across sectors, the drivers of outcome variation continue to be the substantive differences in student populations served and oft correlated variations in access to schooling and non-schooling resources (in public schools, charter schools or private schools).[1]

Why do those KIPP charter middle schools appear to perform so well? What about New York City or Newark charter schools more broadly? And what about years of findings on private schools, or students participating in the New York City private school voucher experiment? It’s not about sectors, but rather about strategies and resources. And if it’s about strategies and resources, then if we can identify what works and the resources needed to legitimately serve all children, we can provide those opportunities within a publicly governed, publicly accountable system of common schools. Indeed, if these measured outcomes were in fact the only issue of concern, we might leverage an appropriately mixed set of schooling providers to get the job done. In fact, the lack of decisive advantage by sector alone is equal justification for agnosticism as it is against it.

But, that’s only if we ignore entirely that there might actually be other tradeoffs involved, beyond whatever test score, graduation, matriculation or employment outcome might be achieved.

Trading Off Legal Rights for Test Scores?

It’s not just about figuring out how to achieve crudely measured “superberific” schooling.  Our children’s schooling exists in a broader social, political and legal context. Kids have legal rights, and under most state constitutions kids a right to access/participate in/gain the benefits of a system of schooling (sometimes, quite explicitly, a system of public schooling). In many states, they not only have a right to access schooling (at times, of some measured degree of quality), but a legal obligation to attend up to a specified age (compulsory schooling laws).

As I’ve discussed on a few previous blog posts, privately governed and/or managed charter schools, more like traditional private than like public schools, may not be (are likely not) subject to the full protection of students’ constitutional or statutory rights (summary table from previous post included below). When attending a private school, it’s clear that kids have no right to continued attendance. They can be expelled, excluded outright for any number of reasons (including admissions testing). They may be compelled to recite school oaths and may be obligated to participate in religious activities and may be restricted in their ability to freely express themselves and subject to disciplinary action including expulsion for failure to comply. Parents may also be obligated to participate in certain activities as a condition of continued enrollment.

While charter advocates love to declare their schools as necessarily “public,” with regard to at least some of these same issues/questions, Charter school legal defense attorneys are quick to argue that they are in fact, private. That, for example, children’s rights under disciplinary codes should be treated as private contracts entered into by parents, just as in private schools – and substantively different from “public” schools – or those formally governed and operated by agents of the state (local elected school boards and public district administrators).

Further, the public-private delineation and murky middle ground of charter schooling raises numerous additional substantive legal questions regarding public employment law and employee rights, taxpayer and citizen rights to open public meetings and public records, and rights, responsibilities, liabilities and protections of “public officials” such as school board members and public employees as opposed to governing boards of private citizens, and employees of private contractors.

Sector agnosticism, as dreadfully simplified by Andrew Smarick requires completely ignoring these substantive tradeoffs.  Trading off constitutional rights to reduce supply of some and increase access to other sectors is not benign, if those sectors could/might possibly yield other advantages.

The Distribution of Lost Rights

Nor do I suspect that the tradeoff of rights will ever be randomly distributed across children by the wealth and income of their families. No-one is asking the superintendent of Scarsdale (great guy, by the way) to Relinquish his schools and adopt a policy of Sector Agnosticism. This is a policy for the children of New Orleans, New York City, Chicago, Philadelphia and Newark.

In the extreme case – a case favored by Smarick and seemingly endorsed (through relinquishment) by Kingsland – a district – or now merely a geographic space – where children have only access to privately governed/managed charter schools may require that any/all that wish to actually exercise their state constitutional right to attend school, have to choose which rights to forgo in the process? Will 100% of parents in that zone be required to enter into contractual agreements (forgoing constitutional & statutory protections) with schools regarding disciplinary policies for their children?

In fact, Kingsland’s logic is that district superintendents should simply succumb or surrender to the forces that wish to forcibly close and takeover their schools, and relinquish those schools or at least the children who would have attended them, to other sectors. Following Smarick’s logic, parents and citizens at large should completely ignore tradeoffs of constitutional protections, or humiliating treatment of children, in lieu of Smarickian measures of “superbification.”

Creating a scenario where only low income minority children in America’s cities must tradeoff their constitutional and statutory protections to gain access to schooling (which they may be compelled to attend) is clearly unacceptable, inequitable treatment.  Before you go there… no… I’m not saying that the responsible policy solution is to make sure that suburban kids and their parents are equally deprived of protections.  What’s not good for some is not good for all.

One logical retort to my arguments here is that if parents want these choices, if they are backed up on waiting lists for existing charters, then we should provide to them. If the demand is there, let the supply meet the demand!? It would be one thing if it was made clear, up front, to potential choosers these hidden tradeoffs, but that’s not the case. If anything, charter advocates are doing their best to conceal that any such tradeoffs exist.

Indeed, appropriate cross-sector regulation might negate some of my concerns raised here, but these issues are too frequently ignored.

Market Manipulation & The Forcible Reduction of the “Public Option”

Worse, in the current policy context, we are not witnessing the emergence of a true, fair and equitable, demand driven and fully open and accessible (driven by open information) system of choice. Policies of relinquishment and sector agnosticism are being pursued in practice as policies of forced relinquishment (read mass closings) of traditional public schooling and sector favoring transfer of assets (public to privately governed charters), coupled with gross misrepresentations of information on sector quality.

In selected cases, we are also witnessing a coordinated effort to provide competitive advantage for non-district alternatives. Where sectors are set up to compete with one another to prove their worth, the likelihood that charter or voucher advocates will lobby for increased resources for district schools is about as likely as the New York Yankee ownership arguing for revenue sharing to help the Kansas City Royals, or Walmart to lobby for tax breaks for Target. Similarly, the likelihood that well endowed charters will share their philanthropy with others less fortunate is slim to none where the emphasis remains on flaunting one’s competitive advantage.

A veneer of demand (as measured by duplicative waiting lists) for private and charter sectors has been induced by forcible reduction of supply of urban schooling, and gross misrepresentations & mismeasures (New Jersey/ New York) of neighborhood schooling quality and manipulation of the playing field.

Closing Thoughts

Before we jump on these reformy bandwagons, and start waiving the white flag of relinquishment and promoting the virtues of sector agnosticism, we need to take a hard look at how this is playing out in our cities. Numerous New Orleans schools were wiped out by a natural disaster, displacing large shares of the lowest income residents to Houston (and elsewhere), many of whom have not been able to return in part because the market based model of New Orleans has chosen not to serve their former, blighted neighborhoods.  This was tragic, and the initial occurrence largely beyond policymakers’ control.  The choice to leave children and their families unserved or underserved was a conscious policy decision (or a least a predictable result of the policy response).

Proposed Chicago (and Philadelphia) school closings would appear comparably poised to induce increased demand for charters, which will likely be used as rationale for expanding charters even further and advancing the cycle toward its ultimate end (as if Katrina by design, and more surgically targeted at schools with low test scores and poor minority children). In most U.S. cities, however charter market shares remain modest, and publicly subsidized private school enrollment even smaller, providing an opportunity to pause and rethink current strategies.

So then, what do we do about all of this? First, reformers and non-reformers alike (and anti-reformers too!) need to step back from these oversimplified talking points and buzz phrases which so illustrate the worst of intellectually lazy, undisciplined, under-informed policy development. I don’t mean to be a hypercritical, ivory tower (actually, public university 1960s era building basement) academic … okay… yeah… that is what I mean to be here. Why? Because it matters! Exploring and understanding these tradeoffs matters. Ignoring them is reckless.


[1] Elite charter schools commonly spend 30 to 50% more than district schools in the same city while often serving much less needy students, and independent private day schools spend nearly double the average of public districts (1.96x) in their same labor market while serving far more advantaged populations.

Supplementary Tables

Governance Issues in LEA and Charter Schooling

Slide1

Governance issues in Voucher and Tuition Tax Credit Programs

Slide2

When Real Life Exceeds Satire: Comments on ShankerBlog’s April Fools Post

Yesterday, Matt Di          Carlo over at Shankerblog put out his April fools post. The genius of the post is in its subtlety.  Matt put together a few graphs of longitudinal NAEP data showing that Maryland had made greater than average national gains on NAEP and then asserted that these gains must therefore be a function of some policy conditions that exist in Maryland. In the Post-RTTT era, Maryland has been the scorn of “reformers” because it just won’t get on board with large scale vouchers and charter expansion and has resisted follow through on test-score based teacher evaluation. Taking a poke a reformy logic, Matt asserted that perhaps the low charter share and lack of emphasis on test score based teacher evaluation… along with a dose of decent funding might be the cause of Maryland’s miracle!

Of course, these assertions are no more a stretch than commonly touted miracles in Texas in the 1990s, Florida or Washington DC, most of which are derived from making loose connections between NAEP trend data and selective discussion of preferred policies that may have concurrently existed.  The difference is that Matt was poking fun at the idea of making bold, decisive, causal inferences from such data. Such data raise interesting questions.

What I found so fun and at the same time deeply disturbing about Matt’s post is that the assertions he made in satire… were nowhere near as absurd as many of the assertions made in studies/reports, etc. I discussed here on my blog over the years. Here are but a few examples of “stuff” presented as serious/legit policy evidence, that make Matt’s satirical assertions seem completely reasonable.

The Many Variations of Money Doesn’t Matter Graphs:

I start with this one, because there are so many versions of it floating around out there, that come and go over time, and are often used to advance the “money doesn’t matter”… we’ve spent ourselves into bankruptcy and gotten nothing for it… graph. Every good reformer has a laminated copy of one version or another of this graph which they carry in wallet-size.

I blogged about this graph when Bill Gates used it in a HuffPo article.

Slide1

Gates asserted:

 Over the last four decades, the per-student cost of running our K-12 schools has more than doubled, while our student achievement has remained flat, and other countries have raced ahead. The same pattern holds for higher education. Spending has climbed, but our percentage of college graduates has dropped compared to other countries… For more than 30 years, spending has risen while performance stayed flat. Now we need to raise performance without spending a lot more.

Among other things, the chart includes no international comparison, which becomes the centerpiece of the policy argument. Beyond that, the chart provides no real evidence of a lack of connection between spending and outcomes across districts within U.S. States.  Instead, the chart juxtaposes completely different measures on completely different scales to make it look like one number is rising dramatically while  the others are staying flat. This tells us NOTHING. It’s just embarrassing. Simply from a graphing standpoint, a blogger at Junk Charts noted:

Using double axes earns justified heckles but using two gridlines is a scandal!  A scatter plot is the default for this type of data. (See next section for why this particular set of data is not informative anyway.)

Not much else to say about that one. Again, had I used an example this absurd to represent reformy research and thinking, I’ d have likely faced stern criticism for mis-characterizing the rigor of reformy research!

This alternate version comes to us from none other than Andrew Coulson of Cato Institute. Coulson has a stellar record of this kind of stuff. So, what would you do to the Gates graph above if you really wanted to make your case that spending has risen dramatically and we’ve gotten no outcome improvement? First, use total rather than per pupil spending (and call it “cost”) and then stretch the scale on the vertical axis for the spending data to make it look even steeper. And then express the achievement data in percent change terms because NAEP scale scores are in the 215 to 220 range for 4th grade reading, for example, but are scaled such that even small point gains may be important/relevant but won’t even show as a blip if expressed as a percent over the base year.

Slide2

Chris Cerf’s Poverty Doesn’t Matter Graph!

Now, it’s one thing when and under-informed tech CEO goes all TED-style on us with big screens, gadgets, bells and whistles and info-graphics that just don’t mean crap anyway. But, it’s yet another when a State Commissioner of Education presents something not only equally ridiculous… but arguably far more ridiculous, disingenuous, unethical and downright WRONG.

This is a graph for the ages, and it comes from a presentation by the New Jersey Commissioner of Education given at the NJASA Commissioner’s Convocation in Jackson, NJ on Feb 29. State of NJ Schools presentation 2-29-2012

Slide4

The title conveys the intended point of the graph – that if you look hard enough across New Jersey – you can find not only some, but MANY higher poverty schools that perform better than lower poverty schools.

This is a bizarre graph to say the least. It’s set up as a scatter plot of proficiency rates with respect to free/reduced lunch rates, but then it only includes those schools/dots that fall in these otherwise unlikely positions. At least put the others there faintly in the background, so we can see where these fit into the overall pattern. The suggestion here is that there is not pattern.

Note: this graph may not even be the worst one in the presentation. You decide!

The apparent inference here? Either poverty itself really isn’t that important a factor in determining student success rates on state assessments, or, alternatively, free and reduced lunch simply isn’t a very good measure of poverty even if poverty is a good predictor. Either way, something’s clearly amiss if we have so many higher poverty schools outperforming lower poverty ones. In fact, the only dots included in the graph are high poverty districts outperforming lower poverty ones. There can’t be much of a pattern between these two variables at all, can there? If anything, the trendline must be sloped uphill? (that is, higher poverty leads to higher outcomes!)

Note that the graph doesn’t even tell us which or how many dots/schools are in each group and/or what percent of all schools these represent. Are they the norm? or the outliers?

Well, here’s what the pattern really looks like with all schools included:

Slide5

Hmmm… looks a little different when you put it that way. Yeah, it’s a scatter, not a perfectly straight line of dots. And yes, there are some dots to the right hand side that land above the 65 line and some dots to the left that land below it.

Note: New Jersey’s Chris Cerf is not alone among state commissioners in promoting completely bogus analysis posing as empirical validation. In fact, New York’s John King presented a completely fabricated graph provided to him by a consultant to the state and has used that graph to frame his state’s policy initiatives.

Rishawn Biddle’s Graph of, well, something? What?

Not to be outdone, Rishawn Biddle who on occasion fashions himself a “researcher” on education policy issues, provides a graph that comes close to the degrees of intentional deception presented by Commissioner Cerf above.  I blogged about this graph here!

In response to arguments I had made on my blog regarding the role of substantive and sustained school finance reforms in improving school quality, Biddle argued:

Despite the arguments (and the pretty charts) of such defenders as Rutgers’ Bruce Baker, there is no evidence that spending more on American public education will lead to better results for children.

My claims are substantiated in this peer reviewed article and this separate more comprehensive report:

  • Baker, B. D., & Welner, K. G. (2011). School Finance and Courts: Does Reform Matter, and How Can We Tell?. Teachers College Record, 113(11), 2374-2414.
  • Baker, B. D. (2012). Revisiting the Age-Old Question: Does Money Matter in Education?. Albert Shanker Institute. http://www.shankerinstitute.org/images/doesmoneymatter_final.pdf

And what does Biddle provide as counter evidence to this – apparent lack of evidence I summarize above (I’ve sent the article link to Biddle on more than one occasion, but he apparently doesn’t read this kind of academic stuff)?

Biddle counters with a link to this graph – a true gem (I’ve added some annotation, not in his original)!

Slide6

Yes, Biddle’s counter to the body of research he has not and likely will never read, is to use this graph of “promoting power” by student race group for Jersey City, NJ in 2004 and 2009. Note that the infusion of additional funds in NJ occurred mainly from 1998 to 2003, leveling off thereafter. But that’s a tangential point (not really).  So, Biddle’s absolute verification that more money doesn’t matter is to simply assert without verification that Jersey City got a whole lot more money and then to use this graph to argue that nothing improved!

First of all, that analysis wouldn’t pass muster in as a master’s degree level assignment (I teach a class on this stuff at that level), no less major research conclusions. From a graphing standpoint, I often criticize my students’ work for what I refer to as gratuitous use of 3d – especially where the use of 3d bars actually obscures the comparisons by making it hard to see where they align on the axis.

But, the really funny if not warped part of this graph is that there appear to be significant gains for black males between 2004 and 2009, but those gains are obscured by hiding the 2009 black male score behind the 2004 black female score.

Note that the graph also contains no information regarding the actual shares of the student population that fall into each group? Not very useful. Pretty damn amateur. Certainly fails to make any particular point, and certainly doesn’t refute the various citations above – all of which employ more rigorous analytic methods, apply to more than a single district, and most of which appear in rigorous peer reviewed journals.

Reason Foundation’s Today’s Policies Affected Yesterday’s Outcomes Study!

Finally, in my years as a reviewer for the National Education Policy Center’s Think Tank Review Project I’ve reviewed a lot of sketchy stuff. Some of it stands out, and has even won Bunkum awards from NEPC.

For example, a recent report from ConnCAN repeatedly footnoted a claim as being substantiated to earlier reports…only to result in a dead end where the claim was never substantiated… and in fact, when checking the data turned out to be patently false!  So, this one isn’t even a subtle data interpretation issue. It’s just a lie.

Then there was a report by the organization Third Way, which gathered numerous sources of incompatible data, across incompatible time frames (along with many other bizarre claims) in order to make the argument that America’s middle class schools are failing miserably.

Either of these reports make Matt’s assertions in his post on the Maryland Miracle look totally reasonable!

But for me, the winner among all of the think tank reports I’ve read comes from the Reason Foundation in their 2009 Weighted Student Funding Yearbook! Here’s the abstract of my review:

The new Weighted Student Formula Yearbook 2009 from the Reason Foundation provides a simple framework for touting the successes of states and urban school districts that grant greater fiscal autonomy to schools. The report defines the Weighted Student Formula (WSF) reform extremely broadly, presenting a variety of reforms under the WSF umbrella. Accordingly, when the report concludes that WSF is successful and should be widely replicated, it is difficult to sort through the claims and recommendations. Moreover, the approach and recommendations lack critical inquiry, thought, or empirical analysis. Perhaps most disturbing is the fact that in a third of the specific districts presented in the report, the evidence of success provided predates the implementation of the reforms, and the Reason press release makes the outright claim that past improvements are somehow a function of yet-to-be-implemented reforms. While the report does provide some reasonable recommendations, they are overshadowed by others. Overall, the policy guidance provided by the Reason report is reckless and irresponsible.

Yes… you read it correctly…. If you go through the smashing successes claimed by Reason in this report, in 1/3 of the cases, the reforms in question were implemented after the window of test scores discussed! Hence, the Bunkum time machine award!

Matt’s satirical example didn’t go anywhere near this far.

In Closing….

In my view, there are at least two lessons from Matt’s post, for either side of the reformy aisle.

First, as I so often point out in my classes on applied data analysis, we need to always take  time to carefully evaluate what our data – whatever data and whatever measures – can and cannot tell us. The latter is key here. Descriptive data can be very useful… as long as we understand what they can and cannot tell us. For that matter, various types of inferential statistical analyses (regression models) can also be useful (and in policy research are often primarily descriptive), but often don’t tell us what we think or would like them to tell us. I’ll likely write more about this topic in the future.

Second, we all should take time to carefully scrutinize the link between empirical evidence and policy assertions (and many should take time to take some legit graduate level research methods and statistics and measurement courses on these topics if they wish to continue to opine so boldly about policy inferences!). Perhaps most importantly we should actually take more time and put more effort into scrutinizing those reports and claims that appear most agreeable to our own predisposed beliefs/opinions.  Everyone has predisposed beliefs (especially those who pretend not to). I would argue that experienced researchers likely have stronger beliefs and opinions… and we should… precisely as a result of years of experience researching specific topics.

Oh… and a third lesson… Don’t make completely BS, false/fabricated/absurd graphs like those above. That’s just ridiculous. Are you kidding me? Hiding 3d bars? (Rishawn?) Deleting most of the cases that define the trend? (Cerf?) That’s just ridiculous! Infuriating! Sickening!

In Connecticut, Where There’s a Reformy Con, There’s a CAN!

I was intrigued a few days ago when I saw this headline in my news alerts regarding school funding.

Headline: Report: Funding helps low-performing school districts

I was particularly intrigued because the headline comes from a Connecticut newspaper where I am fully aware that the state really hasn’t done crap to substantively increase resources for low performing, or more specifically high need schools and districts.

Disclaimer: I am fully aware of this because I have been providing technical/expert assistance to local public school districts that have been persistently shortchanged by the state school finance formula (Education Cost Sharing Formula). That, and even prior to my involvement supporting these districts (and more importantly, the kids they serve) in Connecticut, I had already blogged on their plight.

So then, how can it possibly be that that a CT newspaper would print such a ridiculous headline? And where could one possibly find a “Report” that somehow validates that the state has provided funding to help low performing districts?

Well, in Connecticut, where there’s data-free drivel on education policy spewing from the headlines, there’s usually one single source for that drivel – our old friends at ConnCAN!

Yep, they’ve produced a new report! And it’s about as technically solid as many of their previous reports!

An important caveat here is that the ConnCAN report itself (the linked report) doesn’t really seem to address directly the point that is highlighted in this article – that the reforms being implemented by the Malloy administration have improved the financial conditions of districts serving high need populations.

So then where does this strange assertion come from? Did the author of the “news” (used as loosely as possible) article simply make this up – or were they fed this line by ConnCAN? I’m not sure… but the author of the article in the Middletown newspaper begins with this bold statement:

Funding made available by last year’s Public Act 12-116 has helped some of the states lowest-performing school districts, including Middletown, according to the Connecticut Coalition for Achievement Now, an education advocacy organization based in New Haven.

Then, the author of the article summarizes what are characterized as “Highlights from ConnCAN’s March 2013 Progress Report.”

I find it hard to believe the author of the article crafted these summaries on his/her own. So, let’s take these fact-challenged reformy highlights one at a time (again, on the assumption that these highlights are somehow intended to support the article’s thesis – that the reforms have somehow mitigated funding problems/disparities?):
ConnCAN Con:

School Finance: P.A. 12-116 created a Common Chart of Accounts to be implemented in 2014-15, creating across the board standards aimed at enhancing transparency in education spending. To date, the Office of Policy and Management has selected the accounting firm Blum Shapiro to develop a framework for Common Chart of Accounts development and execution.

MY REPLY

Let’s start here with simple acknowledgement that creating a common chart of accounts does little or nothing – okay, NOTHING – to enhance the equity or adequacy of educational funding across districts. So, what did the state actually do to enhance that funding? Not so much really.

Figure 1 shows the effect of the $50 million dollar increase in ECS Aid for 2012-13, when added to Net Current Expenditures (NCEP) for 2011-12. The 2011-12 NCEP distribution is shown in green dots. The changes to NCEP that would result from the additional state aid are shown in orange dots. In green dots, we see that districts like Bridgeport, New Britain, Waterbury and Meriden are significantly disadvantaged by the ECS formula in 2011-12, in terms of their resultant NCEP.

AND, perhaps more importantly, we see that “increases” to funding for 12-13 really didn’t change much!

Figure 1.

Slide1

Table 1 includes NCEP for 2011-12 and the actual aid increases for 2012-13 (divided by ADM for 11-12) for Alliance Districts which include several high need districts.  I have also expressed the ECS aid increase as a percent increase over NCEP 2011-12. Most increases were less than $200 per pupil and well less than 2%.

Table 1.

Alliance District Spending & Aid Increases 12-13

Malloy_arky
ConnCAN Con:

School Choice: P.A. 12-116 increased per-pupil funding for public charter students ($10,500/FY13, $11,000/FY14, and $11,500/FY15) and allowed for the creation of 4 new state approved charters. Since then, per-pupil charter funds were cut by $300 for the FY13, and 27 letters-of-interest were submitted to the State Department of Education for launching new charters.

MY REPLY:

It is indeed true that recent adjustments to the funding formula provided more significant increases in aid to charter schools.  At best, these increases fail to alter the distribution of opportunities to Connecticut schoolchildren.  More likely, they in fact exacerbate disparities. Charters serve a relatively small share of the total student population. Most children in high need districts remain in district schools that saw negligible increase in funding. In that sense, charter funding increases have limited effect.

But, as it turns out, many of the charter schools in high need districts that received the greater increases in funding actually serve much lower need student populations (See Table 2).

Table 2. Selected Characteristics of Charter Schools in Cities where Mean % Free Lunch Exceeds 50%

Slide2

Further, after removing district expenditures on transportation and special education (expenses for which host districts are primarily responsible), many charters already substantially outspent district averages (see Table 3).[1]

In short, increasing funding to charters which already outspent host districts while cream-skimming lower need students, exacerbates rather than moderating disparities in opportunity.

Table 3. Total & Comparable per Pupil Spending for Charters & Districts with Free/Reduced Lunch >50%, 2009-10, Prior to Funding Boost for Charter Schools

Slide4

[1] Per Pupil Expenditures by Type: http://sdeportal.ct.gov/Cedar/WEB/ct_report/FinanceDTViewer.aspx

[2] Spending on Special Education: http://sdeportal.ct.gov/Cedar/WEB/ct_report/SpecialEducationResourcesDTViewer.aspx

[3] Percent Free or Reduced Lunch: http://sdeportal.ct.gov/Cedar/WEB/ct_report/StudentNeedDTViewer.aspx

ConnCAN Cons (lumping these last two together):

Commissioner’s Network: P.A. 12-116 gave the Commissioner of Education and the State Board of Education authority to select up to 25 of the lowest performing schools into the Commissioner’s Network school turnaround effort. Currently, 4 schools are in the Commissioner’s Network (located in Bridgeport, Hartford, New Haven, and Norwich). The state recently invited six additional schools to submit plans for inclusion in 2013-14 (located in Bridgeport, New Britain, Norwalk, Waterbury (2), and Windham).

Alliance Districts: P.A. 12-116 earmarked $39.5 million in conditional aid for the state’s 30 lowest performing school districts. So far, all 30 district plans have been approved and $39.5 million allocated.

MY REPLY

Now, in the charts above, you’ve seen the rather dramatic (cough/gag) effect that adding $50 million has on Connecticut’s high need districts through the aid formula. Well, here what we have is an even smaller amount of additional aid, to be handed out at the discretion of a single bureaucrat. Nothing systematic. Nothing substantial. Entirely discretionary, and meager.

As noted by ConnCAN, the legislation provides for 25 schools to enter the Commissioner’s network and maybe have access to some additional financial assistance.  There are far more than 25 schools in total in high need districts.  Further, each school can remain in the network for a maximum of three years, and it is unclear whether any supports would exist beyond those three years.

Let’s be absolutely clear here: Educational adequacy and equal educational opportunity a) should not be reserved for a tiny minority of schools, b) should not sunset and c) should not be at the discretion of a single political appointee.

Equally if not more likely, the various proposed structural and governance changes, coupled with new unfunded mandates, will exacerbate existing inequities across Connecticut schools and districts.   For example, many of the policy changes addressed by ConnCAN are little more than labeling schemes that merely highlight existing disparities.

Worse, the most negative and consequential labels fall disproportionately on schools in those districts already disadvantaged financially.

A substantial body of existing literature links school rating systems with local residential property values, including state accountability system assigned school grades.[2]  In short, negative labels may lead to further erosion of housing values and tax base. Further, it is likely that increased threat of state intervention and reduction of local control over schools may adversely affect local property values. The proposed reforms, lacking any substantive provision of additional resources, threaten to accelerate a downward spiral of districts already in long-run economic and educational decline.

Already, a large share of schools classified as “review” schools are not only high need schools, but high need schools concentrated in very high need, and underfunded districts (Bridgeport, Meriden, New Britain, New London & Waterbury).[5]  By contrast, the main distinction of many of the “distinction” schools identified in urban Connecticut contexts is that they serve very few of the lowest income children, few or no children with disabilities and few or no children with limited English language proficiency (See Table 4).

Meanwhile, other schools of distinction are those in the state’s most affluent suburbs.  In other words, the state has adopted a rating scheme driven primarily by student demographics to mislabel the “quality” or “effectiveness” of local public schools. Further, the rating scheme is designed to grant the state greater authority to disrupt local governance of schools, which, while the state may perceive this alternative only in positive light, local property owners and potential property owners may view it quite differently.

Table 4. Selected Characteristics of “Distinction Schools” in Cities where Mean % Free Lunch Exceeds 50%

Slide3ConnCAN Con:

Educator Evaluations: P.A. 12-116 mandated that the educator evaluation program be piloted in 8-10 sites across Connecticut. The Performance Evaluation Advisory Council (PEAC) came to an agreement that the new educator evaluation system would be implemented in all districts with flexibility in 2013-14, and the system would launch statewide with full implementation in 2014-15.

MY REPLY:

Even if one chose to accept that improved teacher evaluation systems and teacher effectiveness measures could be leveraged to better select among teachers on the labor market or in a particular district workforce, our ability to apply that leverage to improve the workforce as a whole, or achieve more equitable distribution of teaching quality would be constrained by a) the overall landscape of teacher compensation relative to other career alternatives and b)  the persistent inequities in financial resources across districts and resulting inequities teacher compensation across advantaged and disadvantaged schools and districts.

The suggestion that mandated changes to teacher evaluation alone will improve the equity and adequacy of the teacher workforce – regardless of resources – ignores that the proposed evaluation models have the potential to significantly increase job uncertainty for teachers without providing increased wages or benefits to counterbalance the risk. Increased job/career and wage expectation uncertainty, while holding wages on average, constant, is likely to lead to reduced, not increased quality of entrants to the profession.

Further, given the emerging body of evidence on the types of metrics proposed for teacher evaluation, career uncertainty is likely to be inequitably distributed, disadvantaging children in already disadvantaged districts and schools.[6]

NOTES


[1] Not accounted for here are potential differences in facilities operation & lease costs. It is often argued that the costs of facilities are particularly high for charter schools, consuming large shares of their budgets, while facilities are “free” for public districts. In reality, one can expect facilities leases for Connecticut charter schools to range from $1,500 per pupil to around $2,000 per pupil (which is indeed significant) and one can expect annual maintenance and operations (not including long term debt expense) for districts to be around $1,400 per pupil (in 2010 based on CTDOE Data).  The state’s choice to provide substantially increased funding for charter schools and not to host district schools was not based on any thorough analysis of actual differences in costs or needs.

[2] Figlio, D. N., & Lucas, M. E. (2004). Whats in a Grade? School Report Cards and the Housing Market. The American Economic Review, 94(3), 591-604.

[6] Baker, B.D., Oluwole, J., Green, P.C. III (2013) The legal consequences of mandating high stakes decisions based on low quality information: Teacher evaluation in the race-to-the-top era. Education Policy Analysis Archives, 21(5). This article is part of EPAA/AAPE’s Special Issue on Value-Added: What America’s Policymakers Need to Know and Understand, Guest Edited by Dr. Audrey Amrein-Beardsley and Assistant Editors Dr. Clarin Collins, Dr. Sarah Polasky, and Ed Sloat. Retrieved [date], from http://epaa.asu.edu/ojs/article/view/1298