Anonymity, Unlinkability, Unobservability, Pseudonymity, and Identity Management – A Consolidated Proposal for Terminology
Andreas Pfitzmann, and Marit Hansen
Abstract
Based on the nomenclature of the early papers in the field, we propose a terminology which is both expressive and precise. More particularly, we define anonymity, unlinkability, unobservability, pseudonymity (pseudonyms and digital pseudonyms, and their attributes), and identity management.
In addition, we describe the relationships between these terms, give a rational why we define them as we do, and sketch the main mechanisms to provide for the properties defined.
http://www.freehaven.net/anonbib/cache/terminology.pdf
Tuesday, March 1, 2011
Leslie Harris: Four Myths You Should Know About 'Do Not Track' Technology.
Diving into 'Do Not Track'
Internet privacy1 is the new black (or at least it's the new green). A recent eye-opening investigative series in the Wall Street Journal exposed the complex web of companies that are "tracking" your every move2 online, and just how much information about you and your online behavior they're collecting in the process.
It's been nearly a decade since Congress first tried and failed to pass a privacy law. That void is a cause of consumer outcry, and Washington is finally listening.
Congress now has several privacy bills — delivered or promised — ready for debate and more are poised for introduction. The Federal Trade Commission and the Commerce Department have both weighed in on how best to approach online consumer privacy. These efforts have converged to create a raucous conversation about how technology and policy can work in concert to protect and enhance Internet privacy3.
Enter Do Not Track4.
What Is Do Not Track?
When you visit a typical commercial website, that website is far from the only one that knows what you have been doing on that site: third-party companies typically contract with websites for permission to track your behavior across many, many sites. Some websites, such as dictionary.com, contract with hundreds of these data collectors (many of whom are advertising networks or are associated with advertising networks).
The information gleaned from your online wanderings is valuable and can be sold to data aggregators who will update your digital dossier or to companies in the advertising industry that will use it to target ads at you. In other words, what happens on dictionary.com -- and most popular websites -- does not stay on dictionary.com.
Many responsible individual advertising networks allow consumers to opt out of this third-party data collection. But consumers cannot be reasonably expected to hunt down every ad network out there and tell them, one at a time, "Do not track me!"
Some self-regulatory initiatives are taking important steps to give consumers more information about behavioral advertising and to make opting out easier and more persistent. For example, the Digital Advertising Alliance is working on technology to put an icon in all targeted ads to let consumers find out why that particular ad showed up in the first place.
The DAA also allows consumers to opt out of all its member companies' serving of behaviorally targeted ads (though not the tracking itself) with just a couple of clicks. Unfortunately, this opt-out is cookie-based, so if you delete your cookies out of concerns for your privacy, you're erasing the instruction to DAA member companies not to track you. (Google recently released a "Keep My Opt Outs" add-on for its Chrome browser that fixes this problem.)
However, even if you do manage to opt out of the behavioral ads of the 60 member companies in the DAA, this doesn't stop the more than 200 other tracking companies out there from following you around the web.
But if web browsers had a Do Not Track feature, then consumers would finally have a simple option to say, "Do not track me!" to every tracking site. The browser companies have recently promised to build is a Do Not Track setting that would allow consumers to prevent companies from tracking them.
Each of the major browser makers -- Google, Microsoft and Mozilla -- is taking a different approach to Do Not Track technology. But Do Not Track is under attack from some who prefer the current balance of power between consumers and trackers to remain tilted in favor of the trackers.
This power struggle has fertilized a number of growing myths on both sides of the debate -- it's time to do some weeding.
Myth: Do Not Track Puts the Government in Charge of the Internet
While , Do Not Track was inspired by the FTC's Do Not Call registry, which has been one of the most successful government-led consumer protection and privacy initiatives in decades, it does not require the same level of government intervention
To make Do Not Track work, the government does not need to maintain any kind of registry. That means no registry of Internet users, no registry of advertisers, no registry of publishers. It simply requires browsers to build the tools and trackers to respect the desires of the consumer.
Government pressure has and should continue to encourage these innovations, but technology platforms don't need the government to get Do Not track off the ground. Many believe that Do Not Track will work only if the government dictates the technology and imposes a single solution on the market, but doing so now will simply stop innovation in its tracks.
The fact is that the major browsers have taken different approaches, and we don't yet know which will work best for consumers. Or whether another approach yet to be unveiled will win the day. We need to encourage a race to the top, not determine a winner prematurely
Myth: We Need a New Law for Do Not Track to Work
It is perhaps overly optimistic to suggest that Do Not Track will work without a "stick" or new legal mechanism for forcing companies to respect consumers' express statement: "Do not track me!"
Legislation that spells out this enforcement mandate has already been proposed, but new legislation is probably not needed to keep companies from violating a Do Not Track request.
The FTC already has the power to pursue companies who engage in deceptive or unfair business practices. The FTC could make clear that violating a Do Not Track request would be considered a deceptive practice and then demonstrate its commitment by bringing enforcement actions.
Myth: Do Not Track Will Kill the Internet
Today, behavioral advertising that links ads to information from tracked consumers makes up a small percentage of the overall multi-billion dollar online ad marketplace. No one knows how many consumers will actually choose to opt out of being tracked, and what's important, Do Not Track does not allow users to opt out of advertising, ad reporting or many types of web analytics.
Users who tell companies "Do not track me!" will see the same number of ads, but these ads will be based on the content of the website they are visiting instead of a dossier of information about the user that has been compiled by tracking their online movements.
Companies also have the option of using Do Not Track as a kind of digital bouncer: If you are blocking their sites, you get no access to their material. Not blocking? You get full website access. That's not a new concept.
Some websites now require users to turn off programs that block ads as a condition to gaining access to their content, much as some websites and videos will load after displaying an advertisement that the consumer cannot avoid. If consumers do not want to be tracked, they can take their business elsewhere.
And some companies could decide to charge for access to their content. A free market is an essential element here: If consumers would rather pay with their wallets instead of paying with their privacy, so be it. If not, that's fine too.
Myth: Do Not Track Is a Panacea for Modern Privacy Problems
It is important not to conflate Do Not Track with comprehensive privacy protection. It is not a silver bullet. It is a solution for a particular privacy concern -- behavioral advertising.
It will not address any of the privacy risks associated with mobile devices, social networking or cloud computing.
Nor will it address the collection of personal data offline or the aggregation and sale of personal data by data brokers. There is a risk that our focus on Do Not Track will divert attention from the real prize: a baseline consumer privacy law that requires all entities that collect and use personal data to engage in fair information practices.
Getting such a law passed is an urgent priority for consumers and for companies that operate in the global economy. Do Not Track is a welcome rest stop along that path, but it must not be the final destination.
Leslie Harris is president and CEO of the Center for Democracy & Technology.
Swire: How Individual Rights Can Both Encourage and Reduce Uses of Personal Information
Social Networks, Privacy, and Freedom of Association
By Peter Swire | February 28, 2011
Governments are concerned about protecting the privacy of social network users and other online activities. Yet a previously unaddressed question is precisely how to create privacy rules without jeopardizing the freedom of association inherent in these networks’ very existence.
The ongoing political transformation in Egypt highlights the crucial role that social networks play in helping individuals organize politically. Facebook was central to the initial sweep of Egyptians onto the streets of their nation’s main cities, allowing dispersed individuals to organize effectively. And democracy protesters could fear, if the popular movement to displace President Hosni Mubarak had not been successful, that the regime would be able to track them down individually, in part through their Facebook accounts.
At precisely the same time that everyday Egyptians were pouring out of their homes in protest, the U.S. Federal Trade Commission was receiving comments on how new online technologies, including social networks, affect privacy. The FTC request obviously did not spark protests across American cities but many here in the United States share the worries of those Egyptian protesters when it comes to privacy, including privacy of their political views but not just political privacy. These deeply held worries about information sharing must be considered given the growing role of social networking in our society—from Barack Obama’s successful online political campaign that helped propel him into the presidency in 2008 to the Tea Party’s successful social networking activism beginning a year later.
This report explores the tension between information sharing, which can promote the freedom of association, and limits on information sharing, notably for privacy protection. Although many experts have written about one or the other, my research has not found any analysis of how the two fit together—how freedom of association interacts with privacy protection. My analysis here, which I offer as a “discussion draft” because the issues have not been explained previously, highlights the profound connection between social networking and freedom of association.
At the most basic level, linguistically, “networks” and “associations” are close synonyms. They both depend on “links” and “relationships.” If there is a tool for lots and lots of networking, then it also is a tool for how we do lots and lots of associations. In this respect, social networks such as Facebook and LinkedIn are simply the latest and strongest associational tools for online group activity, building on email and the Web itself. Indeed, the importance of the Internet to modern political and other group activity is highlighted in a new study by the Pew Foundation, which finds that a majority of online users in the United States have been invited through the Internet to join a group, and a full 38 percent have used the Internet to invite others to join a group.
This new intensity of online associations through social networks is occurring at the same time as social networks and other emerging online activities receive increasing scrutiny from policymakers for privacy reasons, including the Federal Trade Commission, a recent report on privacy from the U.S. Department of Commerce, and a process underway in the European Union to update its Data Protection Directive. All these government efforts are concerned about protecting the privacy of users of social networks and other online activities, yet a previously unaddressed question is precisely how to create privacy rules without jeopardizing the freedom of association inherent in these networks’ very existence.
I stumbled into this tension between association and privacy due to a happenstance of work history. I have long worked and written on privacy and related information technology issues, including as the chief counselor for privacy under President Clinton. Then, during the Obama transition, I was asked to be counsel to the new media team. These were the people who had done such a good job at grassroots organizing during the campaign. During the transition, the team was building new media tools for the transition website and into the overhaul of whitehouse.gov.
My experience historically had been that people on the progressive side of politics often intuitively support privacy protection. They often believe that “they”— meaning big corporations or law enforcement—will grab our personal data and put “us” at risk. The Obama “new media” folks, by contrast, often had a different intuition. They saw personal information as something that “we” use. Modern grassroots organizing seeks to engage interested people and go viral, to galvanize one energetic individual who then gets his or her friends and contacts excited.
In this new media world, “we” the personally motivated use social networks, texts, and other outreach tools to tell our friends and associates about the campaign and remind them to vote. We may reach out to people we don’t know or barely know but who have a shared interest—the same college club, rock band, religious group, or whatever. In this way, “our” energy and commitment can achieve scale and effectiveness. The tools provide “data empowerment,” meaning ordinary people can do things with personal data that only large organizations used to be able to do.
This shift from only “them” using the data to “us” being able to use the data tracks the changes in information technology since the 1970s, when the privacy fair information practices were articulated and the United States passed the Privacy Act. In the 1970s, personal data resided in mainframe computers. These were operated by big government agencies and the largest corporations. Today, by contrast, my personal computer has more processing power than an IBM mainframe from 30 years ago. My home has a fiber-optic connection so bandwidth is rarely a limitation. Today, “we” own mainframes and use the Internet as a global distribution system.
To explain the interaction between privacy and freedom of association, this discussion draft has three sections. The first section explains how privacy debates to date have often featured the “right to privacy” on one side and utilitarian arguments in favor of data use on the other. This section provides more detail about how social networks are major enablers of the right of freedom of association. This means that rules about information flows involve individual rights on both sides, so advocates for either sort of right need to address how to take account of the opposing right.
The second section shows step by step how U.S. law will address the multiple claims of right to privacy and freedom of association. The outcome of litigation will depend on the facts in a particular case but the legal claims arising from freedom of expression appear relevant to a significant range of possible privacy rules that would apply to social networks.
The third section explains how the interesting arguments by New York University law professor Katherine Strandburg fit into the overall analysis. She has written about a somewhat different interaction between privacy and freedom of association, where the right of freedom of association is a limit on the power of government to require an association to reveal its members. As discussed below, her insights are powerful but turn out to address a somewhat different issue than much of the discussion here.
Peter Swire is a Senior Fellow at the Center for American Progress and the C. William O’Neill Professor of Law at the Ohio State University.
Read the full report (pdf)
Download the executive summary (pdf)
Download the report to mobile devices and e-readers from Scribd
By Peter Swire | February 28, 2011
Governments are concerned about protecting the privacy of social network users and other online activities. Yet a previously unaddressed question is precisely how to create privacy rules without jeopardizing the freedom of association inherent in these networks’ very existence.
The ongoing political transformation in Egypt highlights the crucial role that social networks play in helping individuals organize politically. Facebook was central to the initial sweep of Egyptians onto the streets of their nation’s main cities, allowing dispersed individuals to organize effectively. And democracy protesters could fear, if the popular movement to displace President Hosni Mubarak had not been successful, that the regime would be able to track them down individually, in part through their Facebook accounts.
At precisely the same time that everyday Egyptians were pouring out of their homes in protest, the U.S. Federal Trade Commission was receiving comments on how new online technologies, including social networks, affect privacy. The FTC request obviously did not spark protests across American cities but many here in the United States share the worries of those Egyptian protesters when it comes to privacy, including privacy of their political views but not just political privacy. These deeply held worries about information sharing must be considered given the growing role of social networking in our society—from Barack Obama’s successful online political campaign that helped propel him into the presidency in 2008 to the Tea Party’s successful social networking activism beginning a year later.
This report explores the tension between information sharing, which can promote the freedom of association, and limits on information sharing, notably for privacy protection. Although many experts have written about one or the other, my research has not found any analysis of how the two fit together—how freedom of association interacts with privacy protection. My analysis here, which I offer as a “discussion draft” because the issues have not been explained previously, highlights the profound connection between social networking and freedom of association.
At the most basic level, linguistically, “networks” and “associations” are close synonyms. They both depend on “links” and “relationships.” If there is a tool for lots and lots of networking, then it also is a tool for how we do lots and lots of associations. In this respect, social networks such as Facebook and LinkedIn are simply the latest and strongest associational tools for online group activity, building on email and the Web itself. Indeed, the importance of the Internet to modern political and other group activity is highlighted in a new study by the Pew Foundation, which finds that a majority of online users in the United States have been invited through the Internet to join a group, and a full 38 percent have used the Internet to invite others to join a group.
This new intensity of online associations through social networks is occurring at the same time as social networks and other emerging online activities receive increasing scrutiny from policymakers for privacy reasons, including the Federal Trade Commission, a recent report on privacy from the U.S. Department of Commerce, and a process underway in the European Union to update its Data Protection Directive. All these government efforts are concerned about protecting the privacy of users of social networks and other online activities, yet a previously unaddressed question is precisely how to create privacy rules without jeopardizing the freedom of association inherent in these networks’ very existence.
I stumbled into this tension between association and privacy due to a happenstance of work history. I have long worked and written on privacy and related information technology issues, including as the chief counselor for privacy under President Clinton. Then, during the Obama transition, I was asked to be counsel to the new media team. These were the people who had done such a good job at grassroots organizing during the campaign. During the transition, the team was building new media tools for the transition website and into the overhaul of whitehouse.gov.
My experience historically had been that people on the progressive side of politics often intuitively support privacy protection. They often believe that “they”— meaning big corporations or law enforcement—will grab our personal data and put “us” at risk. The Obama “new media” folks, by contrast, often had a different intuition. They saw personal information as something that “we” use. Modern grassroots organizing seeks to engage interested people and go viral, to galvanize one energetic individual who then gets his or her friends and contacts excited.
In this new media world, “we” the personally motivated use social networks, texts, and other outreach tools to tell our friends and associates about the campaign and remind them to vote. We may reach out to people we don’t know or barely know but who have a shared interest—the same college club, rock band, religious group, or whatever. In this way, “our” energy and commitment can achieve scale and effectiveness. The tools provide “data empowerment,” meaning ordinary people can do things with personal data that only large organizations used to be able to do.
This shift from only “them” using the data to “us” being able to use the data tracks the changes in information technology since the 1970s, when the privacy fair information practices were articulated and the United States passed the Privacy Act. In the 1970s, personal data resided in mainframe computers. These were operated by big government agencies and the largest corporations. Today, by contrast, my personal computer has more processing power than an IBM mainframe from 30 years ago. My home has a fiber-optic connection so bandwidth is rarely a limitation. Today, “we” own mainframes and use the Internet as a global distribution system.
To explain the interaction between privacy and freedom of association, this discussion draft has three sections. The first section explains how privacy debates to date have often featured the “right to privacy” on one side and utilitarian arguments in favor of data use on the other. This section provides more detail about how social networks are major enablers of the right of freedom of association. This means that rules about information flows involve individual rights on both sides, so advocates for either sort of right need to address how to take account of the opposing right.
The second section shows step by step how U.S. law will address the multiple claims of right to privacy and freedom of association. The outcome of litigation will depend on the facts in a particular case but the legal claims arising from freedom of expression appear relevant to a significant range of possible privacy rules that would apply to social networks.
The third section explains how the interesting arguments by New York University law professor Katherine Strandburg fit into the overall analysis. She has written about a somewhat different interaction between privacy and freedom of association, where the right of freedom of association is a limit on the power of government to require an association to reveal its members. As discussed below, her insights are powerful but turn out to address a somewhat different issue than much of the discussion here.
Peter Swire is a Senior Fellow at the Center for American Progress and the C. William O’Neill Professor of Law at the Ohio State University.
Read the full report (pdf)
Download the executive summary (pdf)
Download the report to mobile devices and e-readers from Scribd
Saving Facebook - James Grimmelman
ABSTRACT
This Article provides the first comprehensive analysis of the law and policy of privacy on social network sites, using Facebook as its principal example.
It explains how Facebook users socialize on the site, why they misunderstand the risks involved, and how their privacy suffers as a result. Facebook offers a socially compelling platform that also facilitates peer-to-peer privacy violations: users harming each others’ privacy interests. These two facts are inextricably linked; people use Facebook with the goal of sharing information about themselves. Policymakers cannot make Facebook completely safe, but they can help people use it safely.
The Article makes this case by presenting a rich, factually grounded description of the social dynamics of privacy on Facebook. It then uses that description to evaluate a dozen possible policy interventions. Unhelpful interventions—such as mandatory data portability and bans on underage use—fail because they also fail to engage with key aspects of how and why people use social network sites. On the other hand, the potentially helpful interventions—such as a strengthened public-disclosure tort and a right to opt out completely—succeed because they do engage with these social dynamics.
DATA SEGMENTATION IN ELECTRONIC HEALTH INFORMATION EXCHANGE: POLICY CONSIDERATIONS AND ANALYSIS
Prepared for ONC
The issue of whether and, if so, to what extent patients should have control over the sharing or withholding of their health information represents one of the foremost policy challenges related to electronic health information exchange.
It is widely acknowledged that patients’ health information should flow where and when it is needed to support the provision of appropriate and high-quality care. Equally significant, however, is the notion that patients want their needs and preferences to be considered in the determination of what information is shared with other parties, for what purposes, and under what conditions.
Some patients may prefer to withhold or sequester certain elements of health information, often when it is deemed by them (or on their behalf) to be "sensitive," whereas others may feel strongly that all of their health information should be shared under any circumstance.
This discussion raises the issue of data segmentation, which we define for the purposes of this paper as the process of sequestering from capture, access or view certain data elements that are perceived by a legal entity, institution, organization, or individual as being undesirable to share.
This whitepaper explores key components of data segmentation, circumstances for its use, associated benefits and challenges, various applied approaches, and the current legal environment shaping these endeavors.
Full text at: www.gwumc.edu/sphhs/departments/healthpolicy/dhp_publications/pub_uploads/dhpPublication_168F948B-5056-9D20-3D2C53BEED88834B.pdf
It is widely acknowledged that patients’ health information should flow where and when it is needed to support the provision of appropriate and high-quality care. Equally significant, however, is the notion that patients want their needs and preferences to be considered in the determination of what information is shared with other parties, for what purposes, and under what conditions.
Some patients may prefer to withhold or sequester certain elements of health information, often when it is deemed by them (or on their behalf) to be "sensitive," whereas others may feel strongly that all of their health information should be shared under any circumstance.
This discussion raises the issue of data segmentation, which we define for the purposes of this paper as the process of sequestering from capture, access or view certain data elements that are perceived by a legal entity, institution, organization, or individual as being undesirable to share.
This whitepaper explores key components of data segmentation, circumstances for its use, associated benefits and challenges, various applied approaches, and the current legal environment shaping these endeavors.
Full text at: www.gwumc.edu/sphhs/departments/healthpolicy/dhp_publications/pub_uploads/dhpPublication_168F948B-5056-9D20-3D2C53BEED88834B.pdf
Jeff Jonas: How Many Copies of Your Data? Is Somewhat Like Asking: How Many Licks to the Center of the Tootsie Pop?
AUGUST 08, 2007
I get asked from time-to-time how data flows. But, what they really mean is: How many places does the data land? After explaining this a few times I decided to blog it for easy future reference.
If you give a company your name and address, how many copies of this data might there be twelve months later? Many might be surprised to discover that there could easily be in excess of 1,000 copies!
So roughly speaking it looks something like this …
When data first arrives it is likely to be stored in an operational system – sometimes called the "system of record." This is the first instance.
Systems of record are frequently mission critical systems and are therefore candidates for robust backed-up policies. While different organizations have different back-up policies, one common strategy involves creating one backup every day; keeping each daily backup for seven days. This is a rolling strategy where every Monday overwrites last Monday’s backup. An end-of-week backup (e.g., every Sunday night) might be kept for five rolling weeks. Month-end backups might be kept for twelve rolling months. And year-end backups are likely to be kept for something like seven years.
So at the end of twelve months it is possible that there are now an additional 24 copies of the data (7+5+12). The good news is that backups are well protected; the bad news is that the greater the number of backups the greater the chances one turns up missing -- which happens. [Example here]
Structure governs function. [More on this here.] This is important because how the data is structured in the original system of record is specific to its mission. This means if an organization wants to use the data internally for other reasons (e.g., secondary operational systems like a fraud detection system, statistical analysis, marketing, etc.) this data is copied into each additional system.
Along this line, many organizations create a reporting copy that can be used for ad hoc analysis without effecting operational systems. Some copy the data into an operational data store (ODS). Another copy of the data is often moved into to the enterprise data warehouse. Copies from data warehouses are often used to populate data marts. How many data marts might there be? Who knows; one, two, three, or maybe more?
So if an organization has only one reporting copy, one ODS, one enterprise data warehouse and three data marts, then this would add up to six more copies. And these copies are likely to have backups made of them as well, especially when significant computational effort was involved in moving the data (e.g., pre-processed, translation, standardization and integration/co-mingling with secondary data sets). If the same backup strategy is used this could result in 6*24 or 144 more copies.
So now we are at 1 + 24 + 144 = 169 copies.
But wait, there is more. Many of these systems likely have some form of audit logging – maybe both at the application and database level. Often additional "one-time data snapshots" are made over the course of a year for such things as, pre- and post- maintenance and conversion (e.g., application or database upgrades), specialty analysis projects, audit snapshots, and so on. Then there are complete copies made for testing purpose (e.g., to ensure the scheduled upgrade is going to work as planned) and training systems (yes, sometimes training systems are created with real data). These may be backed up as well!
Furthermore, high availability mission critical systems can be expected to have one or more fully synchronized copies of the database strategically dispersed across the landscape for both work load distribution and/or disaster recovery purposes.
And then there are many odd little places data can get parked including sensor-side caching (e.g., at the slot machine or cash register itself), in-transit caches (e.g., cell phone towers), message queues, local and central search engines, performance enhancing indices, and so on.
Sorry, but I’ve lost count. So let’s just say over a hundred copies are made … internally. Now, what about the copies of the data which travel beyond the organization that originally collected the data?
Let’s say you are applying for credit. In this case, you have likely authorized a credit report. Getting your credit report involves sending your information to a (or all three) credit bureau(s). This information request now sits in their system of record; their audit logs; their data warehouses and data marts; their backups and so on. But wait there is more!
These secondary recipients of your data may in turn further disseminate this information. This is especially true if the organization is a data aggregator/data broker. This data is combined with other information, assembled, scored and sold. These tertiary recipients then make their own mission-centric copies, data warehouses, backups, etc. And, in some cases, it is repackaged and sold again.
Care to guess how many copies of the data are out there now?
- No copies You better hope not
- >10 copies Almost certainly
- >100 copies Very likely
- >1,000 copies Quite possible is certain settings
- >100,000 copies Sometimes
- >1,000,000 copies Not out of the question
What can cause information to be replicated over 100,000 times can come into play with such information as phone service (phone books), credit applications, and believe it or even those warranty cards you have been filling out!
What does all this mean?
2. Protecting this many copies of the data is not trivial.
3. The more data you see, the more you realize most data is duplicative.
And this leads to an area I have been thinking about for about five years which I sometimes refer to collectively as "Data Reduction Strategies." More about some progress I have made in this area on some future date.
Oh … and my Perpetual Analytics stuff is going to need one more copy (with its own particular database schema) since Enterprise Intelligence requiresPersistent Context. And, of course, it would be wise to back this up too.
On Search: Metadata - Tim Bray
On Search: Metadata
Tim Bray
In the Web’s early years, the overwhelming favorite among search engines was Yahoo. Today it’s Google. Neither has actually had better text search technology than the competition. They won because they used metadata effectively to make their services more useful. In this ninth On Search episode, a survey of what metadata is, where it comes from, and how to use it. Metadata is technically “information about information” and you can start a fistfight in the bar at any XML or Content-Management conference about what’s data and what’s metadata. In the context of search, metadata is anything that you know about the documents you’re searching beyond the words they contain. With descriptive markup, it’s easy enough to store a document’s metadata right inside it (consider HTML’s tag).
Yahoo · Back when everyone searched at Yahoo, the usual result list looked quite a bit different. If I typed in “donkey,” before the pointers to Web pages there would be a few pointers to categories in the Yahoo taxonomy that contained the word “Donkey.”
This worked really well, because if the Yahoo editor had classified Diseases of the Horse Familyor The Asses of the British Isles under a donkey-related category, I’d find them even though “donkey” wasn’t in the title.
In effect, Yahoo maintained one useful piece of metadata about each page in the engine: What is this about?. This is a real value-add for the searcher.
Google · Google, like Yahoo, maintains one key metadata field about each item it indexes: the well-known PageRank, essentially a measure of how many other pages point to it. They make use of it very simply, to order the result list with high Page-Ranks at the top.
Conclusions? · Google seized search leadership from Yahoo; can we conclude that it’s more important to know how popular something is than to know what it’s about? If you’d told me that ten years ago I would have had a hard time believing it, but the evidence seems pretty compelling. Note that Google actually does have some subject metadata via their integration with the Open Directory Project, but they don’t push it that hard, and the volunteer-staffed, highly-political, AOL-semi-orphan ODP is fairly weak reed to lean on anyhow.
On the other hand, Google has always been way more focused on search than Yahoo has, and isn’t always trying to get in front of you with stock prices and news and weather and so on. More important, even if it turns out that popularity is the key thing for Internet search, the Internet is a very special place, and it’s quite unlikely that popularity is the killer metadatum for the whole universe of search applications.
I believe, though, in the other obvious conclusion: that the number-one way to make search work better is to bring some metadata to bear on the problem. This really shouldn’t be surprising: As I’ve discussed before, it’s really hard to make search engines act much smarter than they do today. So instead, let’s reinforce them with externally-supplied metadata.
Where Does Metadata Come From? · Those Yahoo and Google metadata offerings, while really quite different, have one important thing in common: both are expensive. Yahoo has for years employed a team of editors to sort websites into their subject hierarchy by hand. And Google’s immense rooms full of machines humming away computing PageRanks twenty-four hours a day are a legend in our industry.
In my experience, this is typical. Put another way: There is no cheap metadata. Of course, if we could use computers to compute the metadata like Google does, that would be immensely cheaper than having employees do it. And a lot of smart people have invested a lot of effort and money into the problem of deriving metadata from data, but it’s a hard one. (Still, we should be on the lookout for opportunities; more later).
Many people in the content-management and knowledge-management trades have noticed this, and concluded that the trick is to gather metadata upstream. Remember how Microsoft Word, out of the box, used to pop up a dialog every time you created a new document and encourage you to provide a little metadata? Most people immediately said “Make this go away!” and I don't think Word has done this (by default) for years.
Historically, the difficulty of collecting metadata at source has been generally large enough to outweigh the (potentially huge) benefits from collecting it. But I for one am not ready to give up on this approach. There are, after all, domains where metadata is at the core of the business proposition, and the process works there. For examples, the editorial staff who produce the Wall Street Journal add metadata as they go along, identifying people, companies, stock ticker symbols, and so on.
If You Collect Metadata By Hand · The most important lesson I’ve learned, is: Don’t try to collect too much. You might, just might, get people, when they’re interacting with your intranet, to label their information by project and title; but more than a couple of fields and people will just bypass the process.
This is harder than it looks. When you decide in principle that metadata should be collected, it will develop that many stakeholders have short-lists of the fields they need to make this worthwhile. You can easily end up with a “short” list of a dozen or more fields that constitute the “absolute minimum” that people think you must have. And if you adopt it, you’re dead, because except in special circumstances (e.g. the WSJ), people just will not take the time to do this.
Automatic Metadata · Obviously, there are some metadata items the computer will give you for free: a filename, created/modified dates, who created it, what kind of file (HTML, Excel, PowerPoint), how big it is. These can be handy for search applications and since they’re free, you should collect them and make them available.
The second category of machine-generated metadata is what “Autocategorization” software does. These are the companies like (in alphabetical order) Autonomy, Gavagai, Semio, Stratify, and Vivisimo; they all promise to take your raw data and either generate or fill-in a subject taxonomy telling you what it’s about.
Sometimes they work, sometimes they don’t, and sometimes it can be puzzling figuring out whether they’re going to work or not. But they are not an exception to the no-cheap-metadata rule; this is software that’s generally expensive to buy and expensive to deploy.
Don’t Neglect Your Logfiles · There’s one kind of automatic metadata that I think doesn’t get the respect it deserves: the contents of your logfiles. Here’s the most obvious example: unless you’ve been throwing away your internal Web server log files, you already know which are the most popular items on the Intranet. It would’t be that hard to boil them down (occasionally, on a batch basis, this doesn’t need to be real-time) and develop your own internal “PopRank” based on what gets downloaded the most. It might not be as sexy as PageRank, but if I search the Intranet for material on expense policies, you can bet I’m going to find a lot, and if two or three stand out because they’re the ones everyone ends up reading, you might save a lot of people a lot of time.
Care, Feeding, and Using · Once you’ve got some metadata, since it’s expensive, you should take good care of it. This almost always means putting it in a relational database. As I mentioned above, debates over the meta-ness of data can get religious, but in practice, I’ve observed that while data itself (for example XML or video) often resists being forced into rows and columns, metadata usually lines up happily. Even ongoing has a little MySQL database sitting off to the side of all the XML-encoded entries, tracking a bunch of useful facts about them, including some (e.g. the title) that are replicated inside the data.
And of course you’ll want to put this goodness to work. One obvious way is to have a query screen, so that people can search for resources by author, date, title, and so on, not just brute-force full-text. But what you’d really like is to learn from Yahoo and Google, and have the metadata just there, silently helping. For example, to use in ranking your results.
Another thing you could do is call up Antarctica, our Visual Net product takes metadata and gives search a Graphical User Interface just like your personal computer has.
In the API · This means that if you’re going to design an API for a search engine (something I plan to do eventually in this series) you’re going to need to include entry-points not just for searching and adding words to the full-text index, but also for adding, maintaining, and using the metadata that drives the search.
The Web and the Semantic Web · One of the Web’s distinguishing features is that there’s a big gaping hole where the metadata ought to be. The Web has resources, identified by URI, and you can ask for “representations,” which come with some metadata, but the metadata is about the representation, not the resource. This is probably a bit abstract for those who don’twrestle professionally with Web Architecture, so an example’s in order: Suppose you read an online news story from your desktop computer at 9AM. You get a Web page with some metadata telling you that it’s in HTML and is in English and ISO-Latin-8859-encoded and can’t be cached and so on. Suppose, at noon, on the road, you hit the same story from the minibrowser in your cellphone. The server cleverly notices this is a small-screen device and sends the same information in WAP or simplified HTML or some such thing, with metadata saying what it is (which is completely different from the metadata you got with the PC Browser version).
So, given a URI, the Web has no built-in way to ask questions about it, for example “What is this about?” or “When does it expire?” or “Is this suitable for children?” or “Is this good?”
The Semantic Web project is trying to make the whole Web smarter and more machine-readable, and obviously this is never going to happen without metadata. So a lot of really smart people are working hard to develop good ways to encode, organize, and interchange metadata keyed by URIs. Of course, these people’s dreams aren’t about mere search, they’re about managing your schedule and your medical treatments and your shopping and your supply chain. All of which is fine; but if the Semantic Web ever takes off, there is going to be a whole lot more metadata available about a whole lot of stuff.
As a side-effect, I expect that all the search services of the world will become a lot richer, a lot smarter, and a lot more fun to use. But we’re not there yet.
A Word On Our Sponsor · This is a sponsored essay. It is brought to you by the local power company, who arranged a complete power failure in Antarctica’s offices this afternoon, so I took advantage of battery power to type this in. Power’s back, it’s back to work we go.
Tim Bray
In the Web’s early years, the overwhelming favorite among search engines was Yahoo. Today it’s Google. Neither has actually had better text search technology than the competition. They won because they used metadata effectively to make their services more useful. In this ninth On Search episode, a survey of what metadata is, where it comes from, and how to use it. Metadata is technically “information about information” and you can start a fistfight in the bar at any XML or Content-Management conference about what’s data and what’s metadata. In the context of search, metadata is anything that you know about the documents you’re searching beyond the words they contain. With descriptive markup, it’s easy enough to store a document’s metadata right inside it (consider HTML’s tag).
Yahoo · Back when everyone searched at Yahoo, the usual result list looked quite a bit different. If I typed in “donkey,” before the pointers to Web pages there would be a few pointers to categories in the Yahoo taxonomy that contained the word “Donkey.”
This worked really well, because if the Yahoo editor had classified Diseases of the Horse Familyor The Asses of the British Isles under a donkey-related category, I’d find them even though “donkey” wasn’t in the title.
In effect, Yahoo maintained one useful piece of metadata about each page in the engine: What is this about?. This is a real value-add for the searcher.
Google · Google, like Yahoo, maintains one key metadata field about each item it indexes: the well-known PageRank, essentially a measure of how many other pages point to it. They make use of it very simply, to order the result list with high Page-Ranks at the top.
Conclusions? · Google seized search leadership from Yahoo; can we conclude that it’s more important to know how popular something is than to know what it’s about? If you’d told me that ten years ago I would have had a hard time believing it, but the evidence seems pretty compelling. Note that Google actually does have some subject metadata via their integration with the Open Directory Project, but they don’t push it that hard, and the volunteer-staffed, highly-political, AOL-semi-orphan ODP is fairly weak reed to lean on anyhow.
On the other hand, Google has always been way more focused on search than Yahoo has, and isn’t always trying to get in front of you with stock prices and news and weather and so on. More important, even if it turns out that popularity is the key thing for Internet search, the Internet is a very special place, and it’s quite unlikely that popularity is the killer metadatum for the whole universe of search applications.
I believe, though, in the other obvious conclusion: that the number-one way to make search work better is to bring some metadata to bear on the problem. This really shouldn’t be surprising: As I’ve discussed before, it’s really hard to make search engines act much smarter than they do today. So instead, let’s reinforce them with externally-supplied metadata.
Where Does Metadata Come From? · Those Yahoo and Google metadata offerings, while really quite different, have one important thing in common: both are expensive. Yahoo has for years employed a team of editors to sort websites into their subject hierarchy by hand. And Google’s immense rooms full of machines humming away computing PageRanks twenty-four hours a day are a legend in our industry.
In my experience, this is typical. Put another way: There is no cheap metadata. Of course, if we could use computers to compute the metadata like Google does, that would be immensely cheaper than having employees do it. And a lot of smart people have invested a lot of effort and money into the problem of deriving metadata from data, but it’s a hard one. (Still, we should be on the lookout for opportunities; more later).
Many people in the content-management and knowledge-management trades have noticed this, and concluded that the trick is to gather metadata upstream. Remember how Microsoft Word, out of the box, used to pop up a dialog every time you created a new document and encourage you to provide a little metadata? Most people immediately said “Make this go away!” and I don't think Word has done this (by default) for years.
Historically, the difficulty of collecting metadata at source has been generally large enough to outweigh the (potentially huge) benefits from collecting it. But I for one am not ready to give up on this approach. There are, after all, domains where metadata is at the core of the business proposition, and the process works there. For examples, the editorial staff who produce the Wall Street Journal add metadata as they go along, identifying people, companies, stock ticker symbols, and so on.
If You Collect Metadata By Hand · The most important lesson I’ve learned, is: Don’t try to collect too much. You might, just might, get people, when they’re interacting with your intranet, to label their information by project and title; but more than a couple of fields and people will just bypass the process.
This is harder than it looks. When you decide in principle that metadata should be collected, it will develop that many stakeholders have short-lists of the fields they need to make this worthwhile. You can easily end up with a “short” list of a dozen or more fields that constitute the “absolute minimum” that people think you must have. And if you adopt it, you’re dead, because except in special circumstances (e.g. the WSJ), people just will not take the time to do this.
Automatic Metadata · Obviously, there are some metadata items the computer will give you for free: a filename, created/modified dates, who created it, what kind of file (HTML, Excel, PowerPoint), how big it is. These can be handy for search applications and since they’re free, you should collect them and make them available.
The second category of machine-generated metadata is what “Autocategorization” software does. These are the companies like (in alphabetical order) Autonomy, Gavagai, Semio, Stratify, and Vivisimo; they all promise to take your raw data and either generate or fill-in a subject taxonomy telling you what it’s about.
Sometimes they work, sometimes they don’t, and sometimes it can be puzzling figuring out whether they’re going to work or not. But they are not an exception to the no-cheap-metadata rule; this is software that’s generally expensive to buy and expensive to deploy.
Don’t Neglect Your Logfiles · There’s one kind of automatic metadata that I think doesn’t get the respect it deserves: the contents of your logfiles. Here’s the most obvious example: unless you’ve been throwing away your internal Web server log files, you already know which are the most popular items on the Intranet. It would’t be that hard to boil them down (occasionally, on a batch basis, this doesn’t need to be real-time) and develop your own internal “PopRank” based on what gets downloaded the most. It might not be as sexy as PageRank, but if I search the Intranet for material on expense policies, you can bet I’m going to find a lot, and if two or three stand out because they’re the ones everyone ends up reading, you might save a lot of people a lot of time.
Care, Feeding, and Using · Once you’ve got some metadata, since it’s expensive, you should take good care of it. This almost always means putting it in a relational database. As I mentioned above, debates over the meta-ness of data can get religious, but in practice, I’ve observed that while data itself (for example XML or video) often resists being forced into rows and columns, metadata usually lines up happily. Even ongoing has a little MySQL database sitting off to the side of all the XML-encoded entries, tracking a bunch of useful facts about them, including some (e.g. the title) that are replicated inside the data.
And of course you’ll want to put this goodness to work. One obvious way is to have a query screen, so that people can search for resources by author, date, title, and so on, not just brute-force full-text. But what you’d really like is to learn from Yahoo and Google, and have the metadata just there, silently helping. For example, to use in ranking your results.
Another thing you could do is call up Antarctica, our Visual Net product takes metadata and gives search a Graphical User Interface just like your personal computer has.
In the API · This means that if you’re going to design an API for a search engine (something I plan to do eventually in this series) you’re going to need to include entry-points not just for searching and adding words to the full-text index, but also for adding, maintaining, and using the metadata that drives the search.
The Web and the Semantic Web · One of the Web’s distinguishing features is that there’s a big gaping hole where the metadata ought to be. The Web has resources, identified by URI, and you can ask for “representations,” which come with some metadata, but the metadata is about the representation, not the resource. This is probably a bit abstract for those who don’twrestle professionally with Web Architecture, so an example’s in order: Suppose you read an online news story from your desktop computer at 9AM. You get a Web page with some metadata telling you that it’s in HTML and is in English and ISO-Latin-8859-encoded and can’t be cached and so on. Suppose, at noon, on the road, you hit the same story from the minibrowser in your cellphone. The server cleverly notices this is a small-screen device and sends the same information in WAP or simplified HTML or some such thing, with metadata saying what it is (which is completely different from the metadata you got with the PC Browser version).
So, given a URI, the Web has no built-in way to ask questions about it, for example “What is this about?” or “When does it expire?” or “Is this suitable for children?” or “Is this good?”
The Semantic Web project is trying to make the whole Web smarter and more machine-readable, and obviously this is never going to happen without metadata. So a lot of really smart people are working hard to develop good ways to encode, organize, and interchange metadata keyed by URIs. Of course, these people’s dreams aren’t about mere search, they’re about managing your schedule and your medical treatments and your shopping and your supply chain. All of which is fine; but if the Semantic Web ever takes off, there is going to be a whole lot more metadata available about a whole lot of stuff.
As a side-effect, I expect that all the search services of the world will become a lot richer, a lot smarter, and a lot more fun to use. But we’re not there yet.
A Word On Our Sponsor · This is a sponsored essay. It is brought to you by the local power company, who arranged a complete power failure in Antarctica’s offices this afternoon, so I took advantage of battery power to type this in. Power’s back, it’s back to work we go.
Subscribe to:
Posts (Atom)