By Peter Swire - 06/23/11 07:50 PM ET
More than at any time in the past decade, privacy hearings and proposed legislation are spreading across Capitol Hill. Until now, you could always make money betting against a privacy law passing in Congress. Today, many experts are saying that momentum is building for major legislation, although the shape of that legislation is still unclear.
This round of privacy action is driven by three historic trends, plus other factors that are coming together now.
First is location data. While Apple’s Steve Jobs called the Android a “probe in your pocket,” Apple itself has been brought before both the Senate Judiciary and Commerce committees to try to explain why it was collecting detailed location information on the iPhone. For the first time in history, most Americans are carrying a tracking device — a cell phone — with them in their daily lives. There is great uncertainty about who gets to see that tracking information, including for advertising and law enforcement purposes.
Second is social networking. Facebook has gone from nothing to half a billion users in only a few years. The social networks point out that users voluntarily put that incredible amount of material up on the sites. But this is all so new that the rules of the road are not yet clear.
Third is online behavioral advertising. The Wall Street Journal ran a major series showing the astonishing range of ways that companies can track your activity on the Web — even if you turn off cookies and try to stay anonymous. The companies say that this data is benign, because computers simply choose which ads to show you. Privacy advocates, though, say that these databases give unprecedented insight into what we read and how we think, leading to a scary potential of misuse down the road.
Along with these three mega-trends, Congress is seriously considering federal data-breach legislation, to harmonize state laws and address the Sony PlayStation and other high-profile recent breaches. Major cloud computing companies and civil liberties groups are supporting the Digital Due Process Coalition, which favors a judicial search warrant before law enforcement can gain access to the exabytes of data stored in the cloud. And, there is pressure on the international front, as the European Union considers tightening its own data privacy laws and as India, Mexico and other countries are in the process of putting EU-style privacy laws on the books.
A flashpoint for action could be children’s privacy, where family-values Republicans and consumer-protection Democrats can most easily come together politically. Mark Zuckerberg has publicly discussed bringing under-13s directly into Facebook, but no one knows with what rules. Reps. Edward Markey (D-Mass.) and Joe Barton (R-Texas) have released a discussion draft of the “Do Not Track Kids Act of 2011” to offer the choice not to have behavioral advertising and related tracking for those under the age of 13. And no one knows who will get to see the location information of children — parents will and stalkers won’t, but there are still-to-be-developed rules for those in-between. On June 27, the Center for American Progress will host an event highlighting children’s privacy issues, called “Tracking: Where you are, what you see, and what you do.”
The biggest legislative question might be whether to go with general privacy principles or sector-specific rules. For the first time in history, the administration itself has come out in favor of broad-based privacy legislation for the private sector. The closest fit to the administration vision is the Kerry-McCain “Commercial Privacy Bill of Rights,” which notably would provide individuals with the legal right to opt out of having their information shared for marketing purposes. This sort of general legislation contrasts with sector-specific proposals, such as a recent bill by Sens. Al Franken (D-Minn.) and Richard Blumenthal (D-Conn.) that targets smartphone location information.
With the convergence of all of these technical changes, the current period most resembles the late 1990s. At that time, Congress approved sector-specific laws for medical privacy (HIPAA) and financial services (Gramm-Leach-Bliley), but held off on a general law to protect privacy on the Internet. With so many sectors having specific laws by now, however, the time may well be ripe for a bill that provides basic privacy protections more generally.
Swire was chief counselor for privacy to former President Clinton and served in the National Economic Council under President Obama. He is now a law professor at Ohio State and a fellow with the Center for American Progress and the Future of Privacy Forum.
Friday, June 24, 2011
Thursday, June 16, 2011
Dispelling the Myths Surrounding De-identification: Anonymization Remains a Strong Tool for Protecting Privacy
Dispelling the Myths Surrounding De-identification: Anonymization Remains a Strong Tool for Protecting Privacy
Introduction
Recently, the value of de-identification of personal information as a tool to protect privacy has come into question. Repeated claims have been made regarding the ease of re-identification. We consider this to be most unfortunate because it leaves the mistaken impression that there is no point in attempting to de-identify personal information, especially in cases where de-identified information would be sufficient for subsequent use, as in the case of health research.
The goal of this paper is to dispel this myth — the fear of re-identification is greatly overblown. As long as proper de-identification techniques, combined with re-identification risk measurement procedures, are used, de-identification remains a crucial tool in the protection of privacy. De-identification of personal data may be employed in a manner that simultaneously minimizes the risk of re-identification, while maintaining a high level of data quality. De-identification continues to be a valuable and effective mechanism for protecting personal information, and we urge its ongoing use.
In this paper we illustrate the importance of de-identifying personal information before it is used or disclosed, and at times, prior to its collection. We will demonstrate that, contrary to what has been suggested in recent articles, re-identification of properly de-identified information is not in fact an “easy” or “trivial” task. It requires concerted effort, on the part of skilled technicians. The paper will also describe a tool that minimizes the risk of the re-identification of de-identified information while also enabling a high level of data quality to be maintained. Our objective is to shatter the myth that de-identification is not a strong tool to protect privacy and to ensure that organizations that collect, use and disclose personal information understand the importance of de-identification for the protection of privacy, and continue to use this tool to the greatest extent possible to minimize potential risks. While our primary focus in this paper is on the value of de-identification in the context of personal health information that is used and disclosed for secondary purposes, the same arguments apply in the broader context of personal information.
Introduction
Recently, the value of de-identification of personal information as a tool to protect privacy has come into question. Repeated claims have been made regarding the ease of re-identification. We consider this to be most unfortunate because it leaves the mistaken impression that there is no point in attempting to de-identify personal information, especially in cases where de-identified information would be sufficient for subsequent use, as in the case of health research.
The goal of this paper is to dispel this myth — the fear of re-identification is greatly overblown. As long as proper de-identification techniques, combined with re-identification risk measurement procedures, are used, de-identification remains a crucial tool in the protection of privacy. De-identification of personal data may be employed in a manner that simultaneously minimizes the risk of re-identification, while maintaining a high level of data quality. De-identification continues to be a valuable and effective mechanism for protecting personal information, and we urge its ongoing use.
In this paper we illustrate the importance of de-identifying personal information before it is used or disclosed, and at times, prior to its collection. We will demonstrate that, contrary to what has been suggested in recent articles, re-identification of properly de-identified information is not in fact an “easy” or “trivial” task. It requires concerted effort, on the part of skilled technicians. The paper will also describe a tool that minimizes the risk of the re-identification of de-identified information while also enabling a high level of data quality to be maintained. Our objective is to shatter the myth that de-identification is not a strong tool to protect privacy and to ensure that organizations that collect, use and disclose personal information understand the importance of de-identification for the protection of privacy, and continue to use this tool to the greatest extent possible to minimize potential risks. While our primary focus in this paper is on the value of de-identification in the context of personal health information that is used and disclosed for secondary purposes, the same arguments apply in the broader context of personal information.
Wednesday, June 8, 2011
New Paper: Much Ado About Data Ownership
Much Ado About Data Ownership
Barbara J. Evans
University of Houston Law Center
Harvard Journal of Law and Technology, Vol. 25, Fall 2011
University of Houston Law Center
Abstract:
Recently there have been calls to clarify ownership of data held in large health information networks. This article explores the realities of what patient data ownership would imply to explain why a clearer allocation of entitlements to raw health data would neither enhance patient privacy nor promote access to valuable data resources for public health and research. It updates the debate to account for the 2009 HITECH Act, which correctly recognized that raw patient data are not the valuable resource; these data acquire value only through the application of infrastructure services. The HITECH Act drew on a long tradition of American infrastructure regulation that offers real promise in resolving the infrastructure bottlenecks which (rather than the unresolved status of data ownership) have been the key impediment to data access. Despite this progress there are two unresolved problems, both heretofore neglected in the literature:
First, the existing federal regulatory framework governing data access conceives the state’s police power to use data to promote public health much more narrowly than the police power is conceived in all other legal contexts.
Second, existing regulatory provisions allowing nonconsensual access to data for research fail to incorporate any “public use” requirement to ensure that unconsented research uses of data are justified by a publicly beneficial purpose. As things stand, persons whose health data are used in research have no assurance that the use will serve any socially beneficial purpose at all. This article reframes the debate. The right question is not who owns health data. Instead, the debate should be about appropriate public uses of private data and how best to facilitate them while adequately protecting individuals’ interests.
Download at http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1857986
Barbara J. Evans
University of Houston Law Center
Harvard Journal of Law and Technology, Vol. 25, Fall 2011
University of Houston Law Center
Abstract:
Recently there have been calls to clarify ownership of data held in large health information networks. This article explores the realities of what patient data ownership would imply to explain why a clearer allocation of entitlements to raw health data would neither enhance patient privacy nor promote access to valuable data resources for public health and research. It updates the debate to account for the 2009 HITECH Act, which correctly recognized that raw patient data are not the valuable resource; these data acquire value only through the application of infrastructure services. The HITECH Act drew on a long tradition of American infrastructure regulation that offers real promise in resolving the infrastructure bottlenecks which (rather than the unresolved status of data ownership) have been the key impediment to data access. Despite this progress there are two unresolved problems, both heretofore neglected in the literature:
First, the existing federal regulatory framework governing data access conceives the state’s police power to use data to promote public health much more narrowly than the police power is conceived in all other legal contexts.
Second, existing regulatory provisions allowing nonconsensual access to data for research fail to incorporate any “public use” requirement to ensure that unconsented research uses of data are justified by a publicly beneficial purpose. As things stand, persons whose health data are used in research have no assurance that the use will serve any socially beneficial purpose at all. This article reframes the debate. The right question is not who owns health data. Instead, the debate should be about appropriate public uses of private data and how best to facilitate them while adequately protecting individuals’ interests.
Download at http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1857986
Tuesday, May 31, 2011
NYTimes: Breaches Lead to Push to Protect Medical Data
by Milt Freudenheim • NY Times May 30, 2011
Federal health officials call it the Wall of Shame. It’s a government Web page that lists nearly 300 hospitals, doctors and insurance companies that have reported significant breaches of medical privacy in the last couple of years.
Such lapses, frightening to consumers, could impede the Obama administration’s effort to shift the nation to electronic health care records.
“People need to be assured that their health records are secure and private,” Kathleen Sebelius, secretary of health and human services, said in an interview by phone. “I feel equally strongly that conversion to electronic health records may be one of the most transformative issues in the delivery of health care, lowering medical errors, reducing costs and helping to improve the quality of outcomes.”
So the administration is making new efforts to enforce existing rules about medical privacy and security. But some health care experts wonder if the current rules are enough or whether stronger laws are needed, for example making it a crime for someone to use information obtained improperly.
“The consequences of breaches matter,” conceded Dr. Farzad Mostashari, a former New York public hospitals official who recently became the Obama administration’s national coordinator for health information technology. “People say they are afraid that if their private information becomes known, they may not be able to get health insurance.”
In the last two years, personal medical records of at least 7.8 million people have been improperly exposed, according to the government data. One particularly egregious case involved information about 1.7 million patients, staff members, contractors and suppliers of Bronx hospitals and clinics operated by the Health and Hospitals Corporation, the New York public health agency. Their electronic files were stolen from an unlocked van belonging to a record management company.
The affected patients got the disquieting news that their medical and personal information, like Social Security numbers, had been violated when their health care providers notified them under federal rules.
Showing just how lax security can be, the inspector general of the Department of Health and Human Services said two weeks ago that the agency had found dozens of vulnerabilities in systems to protect records of patients at seven large hospitals in New York, California, Illinois, Texas, Massachusetts, Georgia and Missouri. Auditors cited such problems as personal information that was not encrypted and was stored on computers that could be easily used by unauthorized users.
Auditing teams are now inspecting eight more hospitals, said Lori Pilcher, an assistant inspector general at Health and Human Services. The hospitals are not being identified to avoid alerting hackers, she said.
Another big breach was reported in March on the official Web site by Health Net, a California-based insurance company, which notified 1.9 million health plan members that records with their personal information were missing.
Health Net said I.B.M., which was managing its information system, told the insurer that the records could not be found.
“The health care industry is not as vigilant as they should be about protecting private information in a patient’s medical records,” said Representative Joe L. Barton, a Texas Republican who is co-chairman of the Bipartisan Privacy Caucus in the House.
Mr. Barton knows from personal experience. His own records after a heart attack, along with several thousand others from a research project at the National Institutes of Health, were “on a disk in a laptop in somebody’s trunk that disappeared,” he recalled. “I was stunned.”
The Obama administration has levied a string of stringent penalties for egregious violations of patient rights under the most commonly cited law, the Health Insurance Portability and Accountability Act, or HIPAA, of 1996. Health information is supposed to stay private under those rules, but research has shown that it is not that difficult to connect names and addresses to nominally anonymous data with Internet searches and computerized matchups.
The Office of Civil Rights at Health and Human Services, which took over enforcement of the law, imposed a $1 million fine on Massachusetts General Hospital in March after a hospital employee left paper records of 192 patients on a Boston subway train. The hospital agreed in a settlement, without admitting wrongdoing, to report twice a year on its efforts to tighten patient protections.
Earlier this year, the civil rights office fined a Maryland health plan, Cignet Health, $4.3 million, saying that it had denied patients the right to see their own records in violation of HIPAA provisions. It was the first civil penalty levied under the HIPAA law. “We have ramped up our enforcement,” said Georgina C. Verdugo, director of the civil rights office.
But Dr. David Brailer, a Bush appointee as the first national coordinator of health information technology, is skeptical about whether such efforts will curb security breaches. “We can’t just lock health care data away — because of its role in lifesaving treatment,” Dr. Brailer said.
He said that even with the best technology it would be hard to make health systems secure. “It’s a huge challenge. Break-ins and hacks are unfortunately going to be part of the landscape,” he said.
One protection, he suggested, would be laws to make it illegal for an insurer or employer to discriminate against a person based on information about health conditions like H.I.V./AIDS, cancer and mental health problems.
As a model, he pointed to the antidiscrimination law to prevent the misuse of genetic information that was passed with bipartisan support in the Bush administration. He also said he believed the laws should say “patients own the data, period, and decide what happens to it. The patient should be able to say to Hospital X: ‘send my data to Hospital Y because I’m changing hospitals,’ “ he said.
Today, the information belongs to whoever possesses it, under ideas inherited from 17th-century English common law, he said. “If it gets into your database, essentially you own it,” he added, “and you can pass it on.”
“Today HIPAA makes no sense,” Dr. Brailer added. “The law didn’t anticipate a world where your data passes through many, many hands.”
Wes Rishel, a longtime health care analyst for Gartner, the technology consulting firm, and an adviser to the national coordinator’s office, has a similar view. “Your ability to control access to your information is a horse that is already out of the stable,” he said. “What is really needed is legislation that controls use of it.”
On that score, researchers at Carnegie Mellon University have shown that at least 30 people and organizations have access to the health data of a typical person with private insurance through an employer. They range from pharmacies and drug companies to an employer’s wellness programs and a spouse’s self-insured employer.
“Only you, your doctor and hundreds of others know,” said Latanya Sweeney, a health privacy expert at Harvard and Carnegie Mellon who is also an adviser to the office headed by Dr. Mostashari.
Since HIPAA was enacted there has been “an explosion in data sharing,” Ms. Sweeney said. “And after electronic records are widely adopted, there will be another big explosion.”
Federal health officials call it the Wall of Shame. It’s a government Web page that lists nearly 300 hospitals, doctors and insurance companies that have reported significant breaches of medical privacy in the last couple of years.
Such lapses, frightening to consumers, could impede the Obama administration’s effort to shift the nation to electronic health care records.
“People need to be assured that their health records are secure and private,” Kathleen Sebelius, secretary of health and human services, said in an interview by phone. “I feel equally strongly that conversion to electronic health records may be one of the most transformative issues in the delivery of health care, lowering medical errors, reducing costs and helping to improve the quality of outcomes.”
So the administration is making new efforts to enforce existing rules about medical privacy and security. But some health care experts wonder if the current rules are enough or whether stronger laws are needed, for example making it a crime for someone to use information obtained improperly.
“The consequences of breaches matter,” conceded Dr. Farzad Mostashari, a former New York public hospitals official who recently became the Obama administration’s national coordinator for health information technology. “People say they are afraid that if their private information becomes known, they may not be able to get health insurance.”
In the last two years, personal medical records of at least 7.8 million people have been improperly exposed, according to the government data. One particularly egregious case involved information about 1.7 million patients, staff members, contractors and suppliers of Bronx hospitals and clinics operated by the Health and Hospitals Corporation, the New York public health agency. Their electronic files were stolen from an unlocked van belonging to a record management company.
The affected patients got the disquieting news that their medical and personal information, like Social Security numbers, had been violated when their health care providers notified them under federal rules.
Showing just how lax security can be, the inspector general of the Department of Health and Human Services said two weeks ago that the agency had found dozens of vulnerabilities in systems to protect records of patients at seven large hospitals in New York, California, Illinois, Texas, Massachusetts, Georgia and Missouri. Auditors cited such problems as personal information that was not encrypted and was stored on computers that could be easily used by unauthorized users.
Auditing teams are now inspecting eight more hospitals, said Lori Pilcher, an assistant inspector general at Health and Human Services. The hospitals are not being identified to avoid alerting hackers, she said.
Another big breach was reported in March on the official Web site by Health Net, a California-based insurance company, which notified 1.9 million health plan members that records with their personal information were missing.
Health Net said I.B.M., which was managing its information system, told the insurer that the records could not be found.
“The health care industry is not as vigilant as they should be about protecting private information in a patient’s medical records,” said Representative Joe L. Barton, a Texas Republican who is co-chairman of the Bipartisan Privacy Caucus in the House.
Mr. Barton knows from personal experience. His own records after a heart attack, along with several thousand others from a research project at the National Institutes of Health, were “on a disk in a laptop in somebody’s trunk that disappeared,” he recalled. “I was stunned.”
The Obama administration has levied a string of stringent penalties for egregious violations of patient rights under the most commonly cited law, the Health Insurance Portability and Accountability Act, or HIPAA, of 1996. Health information is supposed to stay private under those rules, but research has shown that it is not that difficult to connect names and addresses to nominally anonymous data with Internet searches and computerized matchups.
The Office of Civil Rights at Health and Human Services, which took over enforcement of the law, imposed a $1 million fine on Massachusetts General Hospital in March after a hospital employee left paper records of 192 patients on a Boston subway train. The hospital agreed in a settlement, without admitting wrongdoing, to report twice a year on its efforts to tighten patient protections.
Earlier this year, the civil rights office fined a Maryland health plan, Cignet Health, $4.3 million, saying that it had denied patients the right to see their own records in violation of HIPAA provisions. It was the first civil penalty levied under the HIPAA law. “We have ramped up our enforcement,” said Georgina C. Verdugo, director of the civil rights office.
But Dr. David Brailer, a Bush appointee as the first national coordinator of health information technology, is skeptical about whether such efforts will curb security breaches. “We can’t just lock health care data away — because of its role in lifesaving treatment,” Dr. Brailer said.
He said that even with the best technology it would be hard to make health systems secure. “It’s a huge challenge. Break-ins and hacks are unfortunately going to be part of the landscape,” he said.
One protection, he suggested, would be laws to make it illegal for an insurer or employer to discriminate against a person based on information about health conditions like H.I.V./AIDS, cancer and mental health problems.
As a model, he pointed to the antidiscrimination law to prevent the misuse of genetic information that was passed with bipartisan support in the Bush administration. He also said he believed the laws should say “patients own the data, period, and decide what happens to it. The patient should be able to say to Hospital X: ‘send my data to Hospital Y because I’m changing hospitals,’ “ he said.
Today, the information belongs to whoever possesses it, under ideas inherited from 17th-century English common law, he said. “If it gets into your database, essentially you own it,” he added, “and you can pass it on.”
“Today HIPAA makes no sense,” Dr. Brailer added. “The law didn’t anticipate a world where your data passes through many, many hands.”
Wes Rishel, a longtime health care analyst for Gartner, the technology consulting firm, and an adviser to the national coordinator’s office, has a similar view. “Your ability to control access to your information is a horse that is already out of the stable,” he said. “What is really needed is legislation that controls use of it.”
On that score, researchers at Carnegie Mellon University have shown that at least 30 people and organizations have access to the health data of a typical person with private insurance through an employer. They range from pharmacies and drug companies to an employer’s wellness programs and a spouse’s self-insured employer.
“Only you, your doctor and hundreds of others know,” said Latanya Sweeney, a health privacy expert at Harvard and Carnegie Mellon who is also an adviser to the office headed by Dr. Mostashari.
Since HIPAA was enacted there has been “an explosion in data sharing,” Ms. Sweeney said. “And after electronic records are widely adopted, there will be another big explosion.”
Thursday, May 26, 2011
A Call for Global Cooperation on Privacy
Live Blogging from the eG8 in Paris: A Call for Global Cooperation on Privacy
My fellow Hogan Lovells Privacy and Information Management practice leader, Marcy Wilder, and I are delegates to the eG8 Forum in Paris, where later today I will be a speaker at the session on privacy…., the gathering has provided a remarkable opportunity for the sharing of ideas and perspectives on the future of the Internet.
Here are my prepared remarks for the privacy session at the eG8 Forum:
My fellow Hogan Lovells Privacy and Information Management practice leader, Marcy Wilder, and I are delegates to the eG8 Forum in Paris, where later today I will be a speaker at the session on privacy…., the gathering has provided a remarkable opportunity for the sharing of ideas and perspectives on the future of the Internet.
Here are my prepared remarks for the privacy session at the eG8 Forum:
- As the only privacy lawyer on today's panel, I appreciate the opportunity to share my perspectives. As we all know, data is the raw material of our Information Age. But the scale and scope of data collection and use are accelerating in ways previously unimaginable. The Internet, mobile devices, and new forms of networked sensors are combining to produce more and more data that can be collected, analyzed, shared and stored. Thus, according to a new McKinsey study we heard about yesterday here at the eG8 Forum, we are entering the era of “big data,” the label for the vast and increasing amounts of digital information being produced every day.
The potential of big data, according to McKinsey, is more efficient and competitive businesses, a stronger world economy and better-served consumers, including with better health care services. The experts at McKinsey are concerned however that before the end of the decade, there will not be enough trained personnel to analyze all of the data.
They also note the issue of personal privacy, an issue underlying the growing concern about the amount of data being collected about our lives and used by businesses, often without our knowledge or consent. While not a focus of the McKinsey study on big data, the world leaders gathering soon in Deauville, France for the annual G8 Summit will be considering the issue of privacy as they address the agenda item on how best to advance the Internet. Presumably, they understand – as a US Commerce Department report recently noted – that if privacy concerns increase, trust in the Internet will decrease, creating an economic drag on the Internet’s potential.
The G8 leaders will be informed by our work. And I hope our discussion of Internet privacy will not divide on geographic lines, with representatives from the EU, which has an omnibus privacy law, expressing disdain for the American targeted approach to privacy protection, and those with a US orientation complaining about over-regulation of privacy. If that is how the discussion evolves, that will be too bad, for there is greater need than ever for global strategies to protect privacy, and countries on both sides of the Atlantic have much to learn from each other.
To be sure, the regional approaches to privacy protection differ even as we share a commitment to the OECD’s Fair Information Practice Principles. In the EU, the Data Protection Directive, implemented through national legislation, is an across-the-board regulation of personal data that places strict limits on the collection, use and retention of personal information. The US, by contrast, has chosen to legislate at the federal level with respect to sensitive data such as health, financial and children’s data, and to target enforcement on privacy violations through the regulatory powers of the Federal Trade Commission and state attorneys general. A number of states have stepped in, too, to regulate the collection, use and security of personal data. Nearly all of the states have data security breach notification laws to inform people when their personal data is at risk.
Privacy self-regulation by businesses and industry groups also is an American tradition, as more and more companies recognize that violations of privacy tarnishes brands and alienates consumers. As the privacy think tank I founded and co-chair, the Future of Privacy Forum, has noted, the recent initiative by industry to empower consumers to stop online tracking of their web activities by advertisers is an example of self-regulatory effort to protect privacy.
While the American approach to privacy may be untidy, in contrast to an omnibus law, a recent Berkeley study concluded that the combination of laws and increased attention by business to the importance of privacy has led to a notably more privacy-protective environment than existed in the 1990’s. And there is recognition in the US that more has to be done to protect privacy. A report from the Federal Trade Commission will be finalized soon on new approaches to privacy protection and legislators on Capitol Hill are focusing on privacy as never before.
Still, the EU takes the position that the US lacks “adequate protection” for the personal data of EU citizens and thus bans the cross-border transfer of such data to the US unless special legal undertakings are made by US businesses to receive the data.
In the US, with our First Amendment traditions, we have trouble understanding the justification for certain EU legal actions in the name of privacy, such as "super injunctions" preventing "tweets" naming litigants in civil actions, enforcement of the so-called “right to be forgotten” against a search engine merely for linking to an unflattering article about someone on the Web. Nor do we understand how a Google executive can be convicted criminally for a random posting by a YouTube user that was said to violate personal privacy.
Despite these differences, there is an emerging consensus on both sides of the Atlantic that people are entitled to greater privacy protections. There is much that can be done cooperatively to advance such protections, like cooperation in cross-border enforcement against multi-national privacy violators, and the adoption of “Privacy by Design” as a standard to be followed by businesses at every stage in the development of new technologies.
In the era of big data, privacy is too important to be overshadowed by claims of legal framework superiority. The eG8 and G8 are good places to sound the chord of cooperation in the advancement of personal privacy.
I am pleased to be part of the discussion.
Tuesday, May 24, 2011
Kerry and McCain: A fair privacy Bill of Rights for online users
By Sen. John McCain (R-Ariz) and Sen. John Kerry (D-Mass.) - 05/23/11 06:27 PM ET
During the past few months, more than 250 million Americans received the frightening news that their personal information, collected by many retailers where they shop, was stolen by hackers routinely.
Sixty-one million Americans who own a smartphone were told that their travels and movements are being tracked by companies who service their smartphones and shared with app providers without restriction on how the information was being used. (We don’t want to imply that Apple or Google phones steal information; they don’t, but the apps on the phones collect and use it without sufficient protections or information for consumers.) And 77 million Americans learned that personal information stored in their online gaming systems was lifted by hackers.
Almost every American is vulnerable to the loss, theft or unanticipated use of their information (theft listed alone is too strong), because in this digital age we routinely turn over personal information to online retailers, social networks and other services in growing numbers.
Americans are rightfully concerned and should be. Is the requirement that you provide such information and cede control of it simply the price of doing business in today’s digital economy? It shouldn’t be. That is why we introduced a Commercial Privacy Bill of Rights — to put Americans back in control of their personal information.
Last year, Internet users sent 107 trillion emails, Facebook hosted 600 million users, Twitter hosted 155 million tweets per day, and Americans across the country shared personal data when checking into hotels, shopping for groceries and refilling their cars. In many ways, all this information sharing is good for consumers. When companies collect data and use it with high ethical standards and the full knowledge and participation of their customers, they can generate immense economic activity, innovate and tailor the services they deliver to the clients they serve.
But today the data collectors are setting the rules. Companies can harvest our personal information and keep it for as long as they like. They can use it and sell it without asking permission. You shouldn’t have to be a computer genius to figure out how to opt out of a company’s information sharing policy. In short, these companies, from mobile phone operators to hotels to websites, can do almost whatever they want with our personal information, and we have no legal right to stop them.
That’s why we introduced the The Commercial Privacy Bill of Rights to keep our private data safe by laying down fair information practices for anyone collecting it. Our legislation will ensure that businesses collecting personal information secure that information, tell people why their data is being collected and allow people to have a say in whether they want their information used. If these companies turn around and transfer this information, any agreements they have made to secure the privacy of their consumers’ information would travel along with it. And if someone requests a company to stop using personal information, they finally have the legal power to make that demand.
We also recognize that it’s important to allow for experimentation and flexibility in the implementation of privacy practices. The Commercial Privacy Bill of Rights does that by establishing voluntary safe-harbor programs to allow companies to design their own privacy programs for complying with the law. They could implement protections however they wanted as long as they still achieved privacy protections on par with the standards set out in the law.
The business community is already responding to the concerns of consumers and regulators by recognizing that the time has come to establish these types of consumer-privacy protections. Industries are negotiating among themselves to establish uniform data collection and use practices. Three of the major Internet browser services have already created tools allowing their users to express their preferences regarding their personal information. Many companies are now making massive investments in privacy protection for their own customers — including employing chief privacy officers to ensure that they earn, retain and respect the trust of consumers.
These companies see that it doesn’t just make good business sense to protect customers’ private information. They know it’s the right thing to do, and we want to take that good work and make it common practice for everyone.
Kerry is the chairman of the Senate Commerce Committee’s subcommittee on Communications, Technology and the Internet. McCain is a former chairman of the Senate Commerce Committee.
Source:http://thehill.com/opinion/op-ed/162781-a-fair-privacy-bill-of-rights-for-online-users
During the past few months, more than 250 million Americans received the frightening news that their personal information, collected by many retailers where they shop, was stolen by hackers routinely.
Sixty-one million Americans who own a smartphone were told that their travels and movements are being tracked by companies who service their smartphones and shared with app providers without restriction on how the information was being used. (We don’t want to imply that Apple or Google phones steal information; they don’t, but the apps on the phones collect and use it without sufficient protections or information for consumers.) And 77 million Americans learned that personal information stored in their online gaming systems was lifted by hackers.
Almost every American is vulnerable to the loss, theft or unanticipated use of their information (theft listed alone is too strong), because in this digital age we routinely turn over personal information to online retailers, social networks and other services in growing numbers.
Americans are rightfully concerned and should be. Is the requirement that you provide such information and cede control of it simply the price of doing business in today’s digital economy? It shouldn’t be. That is why we introduced a Commercial Privacy Bill of Rights — to put Americans back in control of their personal information.
Last year, Internet users sent 107 trillion emails, Facebook hosted 600 million users, Twitter hosted 155 million tweets per day, and Americans across the country shared personal data when checking into hotels, shopping for groceries and refilling their cars. In many ways, all this information sharing is good for consumers. When companies collect data and use it with high ethical standards and the full knowledge and participation of their customers, they can generate immense economic activity, innovate and tailor the services they deliver to the clients they serve.
But today the data collectors are setting the rules. Companies can harvest our personal information and keep it for as long as they like. They can use it and sell it without asking permission. You shouldn’t have to be a computer genius to figure out how to opt out of a company’s information sharing policy. In short, these companies, from mobile phone operators to hotels to websites, can do almost whatever they want with our personal information, and we have no legal right to stop them.
That’s why we introduced the The Commercial Privacy Bill of Rights to keep our private data safe by laying down fair information practices for anyone collecting it. Our legislation will ensure that businesses collecting personal information secure that information, tell people why their data is being collected and allow people to have a say in whether they want their information used. If these companies turn around and transfer this information, any agreements they have made to secure the privacy of their consumers’ information would travel along with it. And if someone requests a company to stop using personal information, they finally have the legal power to make that demand.
We also recognize that it’s important to allow for experimentation and flexibility in the implementation of privacy practices. The Commercial Privacy Bill of Rights does that by establishing voluntary safe-harbor programs to allow companies to design their own privacy programs for complying with the law. They could implement protections however they wanted as long as they still achieved privacy protections on par with the standards set out in the law.
The business community is already responding to the concerns of consumers and regulators by recognizing that the time has come to establish these types of consumer-privacy protections. Industries are negotiating among themselves to establish uniform data collection and use practices. Three of the major Internet browser services have already created tools allowing their users to express their preferences regarding their personal information. Many companies are now making massive investments in privacy protection for their own customers — including employing chief privacy officers to ensure that they earn, retain and respect the trust of consumers.
These companies see that it doesn’t just make good business sense to protect customers’ private information. They know it’s the right thing to do, and we want to take that good work and make it common practice for everyone.
Kerry is the chairman of the Senate Commerce Committee’s subcommittee on Communications, Technology and the Internet. McCain is a former chairman of the Senate Commerce Committee.
Source:http://thehill.com/opinion/op-ed/162781-a-fair-privacy-bill-of-rights-for-online-users
Monday, May 23, 2011
Our data, ourselves
What if privacy is keeping us from reaping the real benefits of the infosphere?
By Leon Neyfakh, The Boston Globe May 22, 2011
If you’re obsessive about your health, and you have $100 to spare, the Fitbit is a portable tracking device you can wear on your wrist that logs, in real time, how many calories you’ve burned, how far you’ve walked, how many steps you’ve taken, and how many hours you’ve slept. It generates colorful graphs that chart your lifestyle and lets you measure yourself against other users. Essentially, the Fitbit is a machine that turns your physical life into a precise, analyzable stream of data.
If this sounds appealing — if you’re the kind of person who finds something seductive about the idea of leaving a thick plume of data in your wake as you go about your daily business — you’ll be glad to know that it’s happening to you regardless of whether you own a fancy pedometer. Even if this thought terrifies you, there’s not much you can do: As most of us know by now, we’re all leaving a trail of data behind us, generating 0s and 1s in someone’s ledger every time we look something up online, make a phone call, go to the doctor, pay our taxes, or buy groceries.
Taken together, the information that millions of us are generating about ourselves amounts to a data set of unimaginable size and growing complexity: a vast, swirling cloud of information about all of us and none of us at once, covering everything from the kind of car we drive to the movies we’ve rented on Netflix to the prescription drugs we take.
Who owns the data in that cloud has been the subject of ferocious debate. It’s not all stored in one place, of course — our lives are tracked and documented by a diffuse assortment of entities that includes private companies like Google and Visa, as well as governmental agencies like the IRS, the Department of Education, and the Census Bureau. Up to now, the public conversation on this kind of data has taken the form of an argument about privacy rights, with legal scholars, computer scientists, and others arguing for tighter restrictions on how our data is used by companies and the government, and consumer advocates instructing us on how to prevent our information from being collected and misused.
But a small group of thinkers is suggesting an entirely new way of understanding our relationship with the data we generate. Instead of arguing about ownership and the right to privacy, they say, we should be imagining data as a public resource: a bountiful trove of information about our society which, if properly managed and cared for, can help us set better policy, more effectively run our institutions, promote public health, and generally give us a more accurate understanding of who we are. This growing pool of data should be public and anonymous, they say — and each of us should feel a civic responsibility to contribute to it.
In a paper forthcoming in the Harvard Journal of Law & Technology, Brooklyn Law School professor Jane Yakowitz introduces the concept of a “data commons” — a sort of public garden where everyone brings their data to be anonymized and made available to researchers working in the public interest. In the paper, she argues that the societal benefits of a thriving data commons outweigh the potential risks from the crooks and hackers who might use it for harm.
Yakowitz’s paper has found support among a wider movement of thinkers who believe that, while protecting people’s privacy is certainly important, it should not be our only priority when it comes to managing information. This position might be a hard sell at a time when consumers are increasingly worried about mass data leaks and identity theft, but Yakowitz and others argue that we shouldn’t let fear of such inevitable accidents cloud our ability to see just how necessary data collection is to our progress as a society.
“There are patterns and trends that none of us can discern by looking at our own individual experiences,” Yakowitz said. “But if we pooled our information, then these patterns can emerge very quickly and irrefutably. So, we should want that sort of knowledge to be made publicly available.”
The idea of sharing one’s personal information with researchers and policy makers for the good of society has a long history in the United States, dating back to the early years of the national census in the 1790s. Back then, a failure to comply with the census was considered a serious abdication of one’s duty to the state. According to Douglas Sylvester, a law professor at Arizona State University, that attitude was grounded in a fundamental belief that in order to run a fair democracy, the country’s leaders needed a detailed knowledge of the people they were governing. Anyone who stood in the way of that was publicly shamed.
“During the early years of the census, your name and your economic information were posted — literally posted, on a sheet of paper — in the public square, for anyone to come and see,” said Sylvester, who has written extensively on the history of data-collection and privacy in America. “The idea was that if your name did not appear, your peers would know that you had not cooperated. Providing this information was a civic obligation.”
Of course, census workers still speak of responsible citizenship and good government when they knock on your door and implore you to fill out their forms, and technically, not doing so is still illegal. But the idea that we owe it to our fellow men to share our information with the public is long gone — and the fact that we think of it as “our” information provides a hint as to what has changed. At some point, privacy experts say, Americans started thinking of their personal data as a form of property, something that could be stolen from them if they didn’t vigilantly protect it.
It’s hard to pinpoint exactly when this transformation began, but its roots lie in the dramatic expansion of administrative data-collection that began around the turn of the last century. A more urban and industrialized nation with more public programs meant that more information was being submitted to government agencies, and eventually, people started getting possessive. Then, during the 1960s, according to Sylvester, the Watergate scandal and advancements in computing power made people even more nervous about government monitoring, and the notion that one’s personal information required protection from hostile outside forces became deeply ingrained in the nation’s psyche.
“Property rules are where people end up going when something is new and uncertain,” said Yakowitz. “When we aren’t sure what to do with something new, there are always a lot of stakeholders who claim a property interest in it. And I think that’s sort of what happened with data.”
Yakowitz came face-to-face with this attitude, and realized how severely it might impede scholarship, as a researcher at UCLA four years ago, when she was working on a study on affirmative action and student performance. Trying to obtain the data sets she needed for her work proved to be an immensely frustrating experience, the 31-year-old said: Some of the schools that kept the records she was after were uncooperative, and in one case, individual graduates who had heard about her research objected to having their information included in her analysis despite the fact that it had been scrubbed of anything that personally identified them.
Yakowitz was disturbed by the fact that her research could be thwarted just because a few people didn’t want “their” data being used in ways they hadn’t anticipated or agreed to. The experience had a galvanizing effect on Yakowitz, causing her to think more pointedly about how Americans understood their relationship to data, and how their attitudes might be at odds with the public interest. Her concept of a “data commons” came out of that thought process. The underlying goal is to revive the idea that sharing our information — this time, without our names attached — should be seen as a civic duty, like a tax that each of us pays in exchange for enjoying the benefits of what researchers learn from it.
Yakowitz began giving presentations on the data commons in February — she visited Google earlier this month to discuss the idea — and although it won’t officially be published until the fall, her paper has already begun attracting attention among people who care about data and privacy law. In it, she reviews the literature on so-called re-identification techniques — the ways that hackers and criminals might cross-reference big, anonymous data sets to figure out information about specific individuals. Yakowitz concludes that these risks have been overblown, and don’t outweigh the social benefits of having lots of anonymized data publicly available. The importance currently placed on privacy in our culture, she says, is the result of a “moral panic” that ultimately hurts us.
She joins a small chorus of voices from law, technology, and government — united under the banner of a movement known as open data — who are already arguing that the benefits of opening up government records and generally disseminating as much data as possible outweigh the costs.
“If you look at the kinds of concerns that we have as a society, they involve questions about health and our economy, and these are all issues which, if they’re to be addressed from an empirical point of view, require actual data on individuals and organizations,” said George T. Duncan, a professor emeritus at Carnegie Mellon University’s Heinz College, who has written about the tension between privacy and the social benefits of data. “Privacy advocates are so locked into their own ideological viewpoint...that they fail to appreciate the value of the data.”
The potential value of data has arguably never been greater, for the simple reason that there’s never been as much of it collected as there is today. According to a report published this month by the consulting firm McKinsey & Co., 15 out of 17 sectors of the American economy have more data stored per company than the entire Library of Congress. One example of data being leveraged for the public good in a way that would have been unthinkable a short time ago is Google Flu Trends, a tool that helps users track the spread of flu by telling them where, and how often, people are typing in flu-related search terms. The Global Viral Forecasting Initiative, based in San Francisco, uses large data sets provided by cellphone and credit card companies to detect and predict epidemics around the world. In Boston recently, a group of researchers commissioned by the governor to study local greenhouse emissions obtained data from the Registry for Motor Vehicles — which keeps inspection records on every car in the city — to find out how much Bostonians were driving.
But advocates of the open data movement see these applications as just a hint of its potential: The more access researchers have to the vast amount of data that is being generated every day, the more accurate and wide-ranging the insights they’ll be able to produce about how to organize our cities, educate our children, fight crime, and stay healthy.
Marc Rodwin, a professor at Suffolk University Law School, has argued for a system in which patient records collected by hospitals and insurance companies — which are currently considered private property, and are routinely purchased in aggregate by pharmaceutical companies — are managed by a central authority and made available, in anonymized form, to researchers. “You can find out about dangerous drugs, you can find out about trends, you can compare effectiveness of different therapies, and the like,” he said. “But if you don’t have that database, you can’t do it.”
Even as such ideas ripen in some corners of the academy and government, proponents of open data are the first to admit that the culture as a whole seems to be heading in the opposite direction. More and more, people are bristling as they realize that everything they do online — including the e-mails they send their friends and the words they search for on Google — is being tracked and turned into data for the benefit of advertisers. And they are made understandably nervous by large-scale data breaches like the one reported last week in Massachusetts, which resulted in as many as 210,000 residents having their financial information exposed to computer hackers. In light of such perceived threats, it’s no wonder the words of privacy advocates are resonating.
Yakowitz and the open data advocates acknowledge that these are reasonable fears, but point out that they won’t be solved by locking down data further. The most damaging breaches, they argue, happen when thieves hack into private sources like credit card processors that are supposedly secure. When we respond by imposing tighter controls on the dissemination of anonymized data, we’re just ensuring that it can’t be used where it might do the most public good.
“The same groups that get really concerned about privacy issues are also the groups that call for more efficiently targeted government resources,” said Holly St. Clair, the director of data services at the Metropolitan Area Planning Council in Boston, where she works on procuring governmental data sets for research purposes. “The only way to do that is with more information — with better information, with more timely information.”
The problem with this vision of the future, according to some privacy experts, is not that large amounts of data don’t come with obvious public benefits. It’s that Yakowitz’s argument presumes a level of anonymization that not only doesn’t exist, but never will. Given enough outside information to draw on, they say, bad actors will always be able to cross-reference data sets with each other, figure out who’s who, and harm individuals who never explicitly agreed to be included in the first place.
In one famous case back in 1997, Carnegie Mellon professor of computer science Latanya Sweeney was able to match publicly available voter rolls to a set of supposedly anonymized medical data, and successfully identify former Massachusetts Governor William F. Weld.
According to Sweeney, currently a visiting professor at Harvard and an affiliate of the Berkman Center for Internet and Society, 87 percent of the US population can be identified by name in this way, based only on birthday, ZIP code, and gender. Sweeney called Yakowitz’s paper on the data commons “irresponsible” for dismissing the risk of re-identification.
There are other practical obstacles as well: Data, in today’s economy, is extremely valuable. Even if data sets could be made truly anonymous, Sweeney asks, why should we expect the huge private collectors of data — companies like Google and Facebook, whose business rather depends on their ability to maintain an exclusive trove of data on their customers — to share what they have for the public good? As data-gathering becomes bigger and bigger business, it might become more valuable to society — but also becomes an asset that companies will fight harder to protect.
As far as Yakowitz is concerned, that’s all the more reason to try to bring about a shift in the way our culture views data. To that end, she proposes granting legal immunity to any entity that releases data into the commons, protecting them from privacy litigation under the condition that they follow a set of strictly enforced standards for anonymization. She also hopes that framing data as a public resource — something that belongs, collectively, to all of us who generate it — will give the public some leverage over big private companies to make their information public.
“Right now I feel like the public gets the rawest deal, because a lot of data is collected, and it’s shared with any company that the private data-collector cares to share it with. But there’s no guarantee that they’ll share it with researchers who are working in the public interest,” Yakowitz said. “Maybe I don’t go far enough — maybe we should force these companies to share with researchers. But that’s for another day, I guess.”
Leon Neyfakh is the staff writer for Ideas. E-mail lneyfakh@globe.com.
By Leon Neyfakh, The Boston Globe May 22, 2011
If you’re obsessive about your health, and you have $100 to spare, the Fitbit is a portable tracking device you can wear on your wrist that logs, in real time, how many calories you’ve burned, how far you’ve walked, how many steps you’ve taken, and how many hours you’ve slept. It generates colorful graphs that chart your lifestyle and lets you measure yourself against other users. Essentially, the Fitbit is a machine that turns your physical life into a precise, analyzable stream of data.If this sounds appealing — if you’re the kind of person who finds something seductive about the idea of leaving a thick plume of data in your wake as you go about your daily business — you’ll be glad to know that it’s happening to you regardless of whether you own a fancy pedometer. Even if this thought terrifies you, there’s not much you can do: As most of us know by now, we’re all leaving a trail of data behind us, generating 0s and 1s in someone’s ledger every time we look something up online, make a phone call, go to the doctor, pay our taxes, or buy groceries.
Taken together, the information that millions of us are generating about ourselves amounts to a data set of unimaginable size and growing complexity: a vast, swirling cloud of information about all of us and none of us at once, covering everything from the kind of car we drive to the movies we’ve rented on Netflix to the prescription drugs we take.
Who owns the data in that cloud has been the subject of ferocious debate. It’s not all stored in one place, of course — our lives are tracked and documented by a diffuse assortment of entities that includes private companies like Google and Visa, as well as governmental agencies like the IRS, the Department of Education, and the Census Bureau. Up to now, the public conversation on this kind of data has taken the form of an argument about privacy rights, with legal scholars, computer scientists, and others arguing for tighter restrictions on how our data is used by companies and the government, and consumer advocates instructing us on how to prevent our information from being collected and misused.
But a small group of thinkers is suggesting an entirely new way of understanding our relationship with the data we generate. Instead of arguing about ownership and the right to privacy, they say, we should be imagining data as a public resource: a bountiful trove of information about our society which, if properly managed and cared for, can help us set better policy, more effectively run our institutions, promote public health, and generally give us a more accurate understanding of who we are. This growing pool of data should be public and anonymous, they say — and each of us should feel a civic responsibility to contribute to it.
In a paper forthcoming in the Harvard Journal of Law & Technology, Brooklyn Law School professor Jane Yakowitz introduces the concept of a “data commons” — a sort of public garden where everyone brings their data to be anonymized and made available to researchers working in the public interest. In the paper, she argues that the societal benefits of a thriving data commons outweigh the potential risks from the crooks and hackers who might use it for harm.
Yakowitz’s paper has found support among a wider movement of thinkers who believe that, while protecting people’s privacy is certainly important, it should not be our only priority when it comes to managing information. This position might be a hard sell at a time when consumers are increasingly worried about mass data leaks and identity theft, but Yakowitz and others argue that we shouldn’t let fear of such inevitable accidents cloud our ability to see just how necessary data collection is to our progress as a society.
“There are patterns and trends that none of us can discern by looking at our own individual experiences,” Yakowitz said. “But if we pooled our information, then these patterns can emerge very quickly and irrefutably. So, we should want that sort of knowledge to be made publicly available.”
The idea of sharing one’s personal information with researchers and policy makers for the good of society has a long history in the United States, dating back to the early years of the national census in the 1790s. Back then, a failure to comply with the census was considered a serious abdication of one’s duty to the state. According to Douglas Sylvester, a law professor at Arizona State University, that attitude was grounded in a fundamental belief that in order to run a fair democracy, the country’s leaders needed a detailed knowledge of the people they were governing. Anyone who stood in the way of that was publicly shamed.
“During the early years of the census, your name and your economic information were posted — literally posted, on a sheet of paper — in the public square, for anyone to come and see,” said Sylvester, who has written extensively on the history of data-collection and privacy in America. “The idea was that if your name did not appear, your peers would know that you had not cooperated. Providing this information was a civic obligation.”
Of course, census workers still speak of responsible citizenship and good government when they knock on your door and implore you to fill out their forms, and technically, not doing so is still illegal. But the idea that we owe it to our fellow men to share our information with the public is long gone — and the fact that we think of it as “our” information provides a hint as to what has changed. At some point, privacy experts say, Americans started thinking of their personal data as a form of property, something that could be stolen from them if they didn’t vigilantly protect it.
It’s hard to pinpoint exactly when this transformation began, but its roots lie in the dramatic expansion of administrative data-collection that began around the turn of the last century. A more urban and industrialized nation with more public programs meant that more information was being submitted to government agencies, and eventually, people started getting possessive. Then, during the 1960s, according to Sylvester, the Watergate scandal and advancements in computing power made people even more nervous about government monitoring, and the notion that one’s personal information required protection from hostile outside forces became deeply ingrained in the nation’s psyche.
“Property rules are where people end up going when something is new and uncertain,” said Yakowitz. “When we aren’t sure what to do with something new, there are always a lot of stakeholders who claim a property interest in it. And I think that’s sort of what happened with data.”
Yakowitz came face-to-face with this attitude, and realized how severely it might impede scholarship, as a researcher at UCLA four years ago, when she was working on a study on affirmative action and student performance. Trying to obtain the data sets she needed for her work proved to be an immensely frustrating experience, the 31-year-old said: Some of the schools that kept the records she was after were uncooperative, and in one case, individual graduates who had heard about her research objected to having their information included in her analysis despite the fact that it had been scrubbed of anything that personally identified them.
Yakowitz was disturbed by the fact that her research could be thwarted just because a few people didn’t want “their” data being used in ways they hadn’t anticipated or agreed to. The experience had a galvanizing effect on Yakowitz, causing her to think more pointedly about how Americans understood their relationship to data, and how their attitudes might be at odds with the public interest. Her concept of a “data commons” came out of that thought process. The underlying goal is to revive the idea that sharing our information — this time, without our names attached — should be seen as a civic duty, like a tax that each of us pays in exchange for enjoying the benefits of what researchers learn from it.
Yakowitz began giving presentations on the data commons in February — she visited Google earlier this month to discuss the idea — and although it won’t officially be published until the fall, her paper has already begun attracting attention among people who care about data and privacy law. In it, she reviews the literature on so-called re-identification techniques — the ways that hackers and criminals might cross-reference big, anonymous data sets to figure out information about specific individuals. Yakowitz concludes that these risks have been overblown, and don’t outweigh the social benefits of having lots of anonymized data publicly available. The importance currently placed on privacy in our culture, she says, is the result of a “moral panic” that ultimately hurts us.
She joins a small chorus of voices from law, technology, and government — united under the banner of a movement known as open data — who are already arguing that the benefits of opening up government records and generally disseminating as much data as possible outweigh the costs.
“If you look at the kinds of concerns that we have as a society, they involve questions about health and our economy, and these are all issues which, if they’re to be addressed from an empirical point of view, require actual data on individuals and organizations,” said George T. Duncan, a professor emeritus at Carnegie Mellon University’s Heinz College, who has written about the tension between privacy and the social benefits of data. “Privacy advocates are so locked into their own ideological viewpoint...that they fail to appreciate the value of the data.”
The potential value of data has arguably never been greater, for the simple reason that there’s never been as much of it collected as there is today. According to a report published this month by the consulting firm McKinsey & Co., 15 out of 17 sectors of the American economy have more data stored per company than the entire Library of Congress. One example of data being leveraged for the public good in a way that would have been unthinkable a short time ago is Google Flu Trends, a tool that helps users track the spread of flu by telling them where, and how often, people are typing in flu-related search terms. The Global Viral Forecasting Initiative, based in San Francisco, uses large data sets provided by cellphone and credit card companies to detect and predict epidemics around the world. In Boston recently, a group of researchers commissioned by the governor to study local greenhouse emissions obtained data from the Registry for Motor Vehicles — which keeps inspection records on every car in the city — to find out how much Bostonians were driving.
But advocates of the open data movement see these applications as just a hint of its potential: The more access researchers have to the vast amount of data that is being generated every day, the more accurate and wide-ranging the insights they’ll be able to produce about how to organize our cities, educate our children, fight crime, and stay healthy.
Marc Rodwin, a professor at Suffolk University Law School, has argued for a system in which patient records collected by hospitals and insurance companies — which are currently considered private property, and are routinely purchased in aggregate by pharmaceutical companies — are managed by a central authority and made available, in anonymized form, to researchers. “You can find out about dangerous drugs, you can find out about trends, you can compare effectiveness of different therapies, and the like,” he said. “But if you don’t have that database, you can’t do it.”
Even as such ideas ripen in some corners of the academy and government, proponents of open data are the first to admit that the culture as a whole seems to be heading in the opposite direction. More and more, people are bristling as they realize that everything they do online — including the e-mails they send their friends and the words they search for on Google — is being tracked and turned into data for the benefit of advertisers. And they are made understandably nervous by large-scale data breaches like the one reported last week in Massachusetts, which resulted in as many as 210,000 residents having their financial information exposed to computer hackers. In light of such perceived threats, it’s no wonder the words of privacy advocates are resonating.
Yakowitz and the open data advocates acknowledge that these are reasonable fears, but point out that they won’t be solved by locking down data further. The most damaging breaches, they argue, happen when thieves hack into private sources like credit card processors that are supposedly secure. When we respond by imposing tighter controls on the dissemination of anonymized data, we’re just ensuring that it can’t be used where it might do the most public good.
“The same groups that get really concerned about privacy issues are also the groups that call for more efficiently targeted government resources,” said Holly St. Clair, the director of data services at the Metropolitan Area Planning Council in Boston, where she works on procuring governmental data sets for research purposes. “The only way to do that is with more information — with better information, with more timely information.”
The problem with this vision of the future, according to some privacy experts, is not that large amounts of data don’t come with obvious public benefits. It’s that Yakowitz’s argument presumes a level of anonymization that not only doesn’t exist, but never will. Given enough outside information to draw on, they say, bad actors will always be able to cross-reference data sets with each other, figure out who’s who, and harm individuals who never explicitly agreed to be included in the first place.
In one famous case back in 1997, Carnegie Mellon professor of computer science Latanya Sweeney was able to match publicly available voter rolls to a set of supposedly anonymized medical data, and successfully identify former Massachusetts Governor William F. Weld.
According to Sweeney, currently a visiting professor at Harvard and an affiliate of the Berkman Center for Internet and Society, 87 percent of the US population can be identified by name in this way, based only on birthday, ZIP code, and gender. Sweeney called Yakowitz’s paper on the data commons “irresponsible” for dismissing the risk of re-identification.
There are other practical obstacles as well: Data, in today’s economy, is extremely valuable. Even if data sets could be made truly anonymous, Sweeney asks, why should we expect the huge private collectors of data — companies like Google and Facebook, whose business rather depends on their ability to maintain an exclusive trove of data on their customers — to share what they have for the public good? As data-gathering becomes bigger and bigger business, it might become more valuable to society — but also becomes an asset that companies will fight harder to protect.
As far as Yakowitz is concerned, that’s all the more reason to try to bring about a shift in the way our culture views data. To that end, she proposes granting legal immunity to any entity that releases data into the commons, protecting them from privacy litigation under the condition that they follow a set of strictly enforced standards for anonymization. She also hopes that framing data as a public resource — something that belongs, collectively, to all of us who generate it — will give the public some leverage over big private companies to make their information public.
“Right now I feel like the public gets the rawest deal, because a lot of data is collected, and it’s shared with any company that the private data-collector cares to share it with. But there’s no guarantee that they’ll share it with researchers who are working in the public interest,” Yakowitz said. “Maybe I don’t go far enough — maybe we should force these companies to share with researchers. But that’s for another day, I guess.”
Leon Neyfakh is the staff writer for Ideas. E-mail lneyfakh@globe.com.
Subscribe to:
Posts (Atom)