County level is nice, but it would be far more efficacious to be able to drill down to zipcode level. Many counties like mine, Hennepin County, MN, is HUGE and has a WIDE RANGE of density.
Data to my understanding is not collected on that level. I haven't seen a source so far that goes to the CENUS tract or ZIP level. I agree that would be really great, just data is not being collected at that spatial scale.
To second what @aaraney is saying, I know that, for San Diego, we have county level data, we have city level data, but we do not have zip code level data.
No one would disagree that having better data would be ideal. As far as I know NYT is not collecting data, only compiling it. Consider closing wishlist items for now so the issues can stay focused on problems within the dataset.
Zip codes don't have their own government bureaucracies to report or account for this, or anything.
Zip codes seen unlikely, but perhaps some of the data is available for city or hospital, and could be compiled here?
The reason the zip code isn’t included is because It is protected health information under HIPAA. Including the zip code would potentially identify some patients. Scientists could look at this data if they had approval from a research ethics board, but could not include that with a public (de-identified) dataset.
Some places, Alberta Canada for example, provide the public with aggregate data by urban administrative districts smaller than county. This is extremely useful for mapping, and seeing what parts of a community are hardest hit, where change is occurring most rapidly. If data collected on each case does not include zip code or some other field smaller than county, hopefully this pandemic experience will make it abundantly clear why it should be collected.
I'm not familiar with the laws in Canada, but for the US, use of the zip code is restricted as follows:
"All geographical subdivisions smaller than a State, including street address, city, county, precinct, zip code, and their equivalent geocodes, except for the initial three digits of a zip code, if according to the current publicly available data from the Bureau of the Census: (1) The geographic unit formed by combining all zip codes with the same three initial digits contains more than 20,000 people; and (2) The initial three digits of a zip code for all such geographic units containing 20,000 or fewer people is changed to 000." (https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html#standard)
Aggregate data is helpful, however, there will still be cases, even with just the first three digits of the zip code, where there are currently so few patients that aggregating their data would potentially identify people.
The health agencies that are receiving the reports of this disease may have access to additional data points for analysis, but are prevented from releasing them to the public because of the law. So these things can be studied, but the dataset must be de-identified if it is to be released publicly.
If you want to go down a rabbit hole, this is a great paper by Paul Ohm describing the various ways that people can be identified even when the dataset is presumed de-identified. While I agree the exact zip code would make the dataset more useful from a scientific standpoint, this would increase the risk of identification to the individuals. As we've seen with other outbreaks, there can be significant stigma and even violence towards those who are identified as infected or associated with a group that was infected. So the protection of privacy is not a theoretical idea.
Shannon,
Yes, I am familiar with and understand that.
But I also mentioned use of other types of divisions. Data could be aggregated by health district, health region, borough, etc., whatever works to provide some sense of geographic distribution smaller than county in areas with fairly high population density. Some US places do this, many do not.
The laws in Canada are not so different in these regards. I included that in case you wanted to go online and see what Alberta makes public and get a sense of how useful it can be.
Stay safe,
-Joan Cunningham
On Mar 30, 2020, at 09:24, ShannonMR notifications@github.com wrote:

I'm not familiar with the laws in Canada, but for the US, use of the zip code is restricted as follows:"All geographical subdivisions smaller than a State, including street address, city, county, precinct, zip code, and their equivalent geocodes, except for the initial three digits of a zip code, if according to the current publicly available data from the Bureau of the Census: (1) The geographic unit formed by combining all zip codes with the same three initial digits contains more than 20,000 people; and (2) The initial three digits of a zip code for all such geographic units containing 20,000 or fewer people is changed to 000." (https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html#standard)
Aggregate data is helpful in this case, however, the zip code would need to be known in order to do this and there will still be areas, even with just the first three digits of the zip code, where there are currently so few patients that aggregating their data wouldn't work.
Remember, this is a public dataset - the health agencies that are receiving the reports of this disease may have access to additional data points for analysis, but are prevented from releasing them to the public. So these things can be studied, but the dataset must be de-identified if it is to be released publicly.
—
You are receiving this because you commented.
Reply to this email directly, view it on GitHub, or unsubscribe.
I should clarify that I don't have any association with the NYT dataset. I saw a number of people commenting on the lack of zip codes and shared my understanding of why the data wasn't presented this way. There is no doubt that aggregate data could (and will) be helpful in understanding this pandemic.
Best,
Shannon
Thanks all for the discussion. We agree that data at ZIP code-level would be great to have. Unfortunately, counties are the smallest geography for which we can consistently get data from across the country.
If zip-code data are available from some regions, and if the regulations would allow its inclusion, that data could be useful for those regions.
-JC
On Mar 30, 2020, at 4:30 PM, Albert Sun notifications@github.com wrote:
Thanks all for the discussion. We agree that data at ZIP code-level would be great to have. Unfortunately, counties are the smallest geography for which we can consistently get data from across the country.
—
You are receiving this because you commented.
Reply to this email directly, view it on GitHub https://github.com/nytimes/covid-19-data/issues/16#issuecomment-606261990, or unsubscribe https://github.com/notifications/unsubscribe-auth/AO7UPQMQT7SZAIFDA5M2PNTRKEFOJANCNFSM4LVFLNHQ.
Most helpful comment
No one would disagree that having better data would be ideal. As far as I know NYT is not collecting data, only compiling it. Consider closing wishlist items for now so the issues can stay focused on problems within the dataset.