{"id":7,"date":"2008-11-20T20:02:45","date_gmt":"2008-11-21T01:02:45","guid":{"rendered":"http:\/\/www.bitquill.net\/blog\/?p=7"},"modified":"2016-05-12T14:51:21","modified_gmt":"2016-05-12T19:51:21","slug":"on-data-ownership-in-a-networked-world","status":"publish","type":"post","link":"http:\/\/bitquill.net\/blog\/on-data-ownership-in-a-networked-world\/","title":{"rendered":"On data ownership in a networked world"},"content":{"rendered":"<p>Every piece of content has a creator and owner (in this post, I will assume they are by default the same entity).  I do not mean ownership in the traditional sense of, e.g., stashing a piece of paper in a drawer, but in the metaphysical sense that each artifact is forever associated with one or more &#8220;creators.&#8221;<\/p>\n<p>This is certainly true of the end-products of intellectual labor, such as the article you are reading.  However, it is also true of more mundane things, such as checkbook register entries or credit card activity. Whenever you pay a bill or purchase an item, you implicitly &#8220;create&#8221; a piece of content: the associated entry in your statement.\u00c2\u00a0 This has two immediately identifiable &#8220;creators&#8221;: the payer (you) and the payee.\u00c2\u00a0 The same is true for, e.g., your email, your IM chats, your web searches, etc. Interesting tidbit: <a title=\"Udi Manber - Search is a Hard Problem\" href=\"http:\/\/www.readwriteweb.com\/archives\/udi_manber_search_is_a_hard_problem.php\">over 20% of search terms entered daily in Google are new<\/a>, which would imply roughly 20 million new pieces of content per day, or over 7 billion (over twice the earth&#8217;s population) per year\u00e2\u20ac\u201dall this from just one activity on one website.<\/p>\n<p>When I spend a few weeks working on, say, a research paper, I have certain expectations and demands about my rights as a &#8220;creator.&#8221; However, I give almost no thought to my rights on the trail of droppings (digital or otherwise) that I &#8220;create&#8221; each day, by searching the web, filling up the gas tank, getting coffee, going through a toll booth, swiping my badge, and so on.\u00c2\u00a0 <strong>However, with the increasing ease of data collection and distribution in digital form, we should re-think our attitudes towards &#8220;authorship&#8221;.<\/strong><\/p>\n<p><!--more--><\/p>\n<h4>Unique identity<\/h4>\n<p>People call me &#8220;Spiros&#8221;, my identity documents list me as &#8220;Spyridon Papadimitriou&#8221; and on most online sites I&#8217;m registered as <tt>spapadim<\/tt>.\u00c2\u00a0 However, sometimes I&#8217;m <tt>s_papadim<\/tt> or <tt>spiros_papadimitriou<\/tt>, and so on.\u00c2\u00a0 Like most people, I lost track of all my accounts a time ago.\u00c2\u00a0 Vice versa, I&#8217;m not the only &#8220;Spiros Papadimitriou&#8221; in the real world.\u00c2\u00a0 For example, I occasionally get confused with my cousin, and receive comments about my interesting architectural designs!\u00c2\u00a0 Nor am I the only <tt>spapadim<\/tt> on the net.<\/p>\n<p><strong>A framework and mechanisms that allow (but do not enforce) asserting and verifying which of those labels (i.e., names, userids, etc) refer to the same entity (i.e., me) is missing.<\/strong> However, this is a prerequisite: how can we talk about data ownership and tackle portability, transparency and accountability, if we have to jump through countless hoops just to prove identity?<\/p>\n<p>Some people, especially in the US, may object or even outright panic at the thought of such a global identifier.\u00c2\u00a0 In Greece, and in much of Europe, we&#8217;ve had national identity cards for decades.\u00c2\u00a0 Which is fine, as long as you know they exist and what are permissible uses-in other words, as long as transparency is ensured.\u00c2\u00a0 Furthermore, the <em>illusion<\/em> of privacy should not be confused with privacy itself\u00e2\u20ac\u201dif in doubt, I suggest reading &#8220;<a title=\"Database Nation (Google Books, no preview)\" href=\"http:\/\/books.google.com\/books?id=T_7TGAAACAAJ\">Database Nation<\/a>&#8221; (<a title=\"Database Nation (official site)\" href=\"http:\/\/www.databasenation.com\/\">official site<\/a>).\u00c2\u00a0 Its examples are largely US-centric, but the lessons are not.<\/p>\n<p><a title=\"OpenID website\" href=\"http:\/\/openid.net\/\">OpenID<\/a> (despite <a title=\"OpenID Phishing Brainstorm (OpenID wiki)\" href=\"http:\/\/wiki.openid.net\/OpenID_Phishing_Brainstorm\">some shortcomings<\/a>) and <a title=\"OAuth website\" href=\"http:\/\/oauth.net\/\">OAuth<\/a> are emerging as open standards for authentication and authorization.\u00c2\u00a0 OpenID allows reuse of authentication credentials from one site on others: I can reuse, say, my Google username and password to log in to other sites (e.g., to leave a comment on this blog), without having to create yet another account from scratch.\u00c2\u00a0 OAuth resembles Kerberos&#8217;s ticket granting service but for the web, permitting other web services to ask for access to a subset of personal information: I could allow Facebook to access only my Google addressbook and not, potentially, all of my data on any Google service.\u00c2\u00a0 OpenID and OAuth can, at least in principle, work together.<\/p>\n<p>Both high-profile individual developers and major companies are involved in these efforts.\u00c2\u00a0 For example, Yahoo! <a title=\"OpenID - Yahoo!\" href=\"http:\/\/openid.yahoo.com\/\">already supports OpenID<\/a> and <a title=\"OAuth - Yahoo! Developer\" href=\"http:\/\/developer.yahoo.com\/oauth\/\">plans to support OAuth<\/a> as well, while Google <a title=\"OAuth - Google Accounts API\" href=\"http:\/\/code.google.com\/apis\/accounts\/docs\/OAuth.html\">supports OAuth<\/a> directly and OpenID indirectly in <a title=\"OpenID - Google AppSpot\" href=\"http:\/\/openid-provider.appspot.com\/\">various<\/a> <a title=\"Google Offers OpenID Logins Via Blogger - TechCrunch\" href=\"http:\/\/www.techcrunch.com\/2008\/01\/18\/google-offers-openid-logins-via-blogger\/\">ways<\/a>.\u00c2\u00a0 Wide adoption of these standards would be a major step forwards for data portability and web interoperability.\u00c2\u00a0 However, I suspect they fall slightly short of providing a truly permanent and global personal identity.\u00c2\u00a0 What if, for any reason, my Yahoo! account disappears, either because I decided to shut it down or because Yahoo! went bust?<\/p>\n<p>I was going to suggest a DNS-based solution and I was surprised when I found that the generic top-level domain <a title=\".name - Wikipedia\" href=\"http:\/\/en.wikipedia.org\/wiki\/.name\"><tt>.name<\/tt><\/a> has been instituted since 2001 to provide URIs for personal identities. You can register for a free three-month trial on <a title=\"FreeYourID website\" href=\"http:\/\/www.freeyourid.com\/\">FreeYourID<\/a> (after that, it&#8217;s $11\/year). What&#8217;s more, their service already provides OpenID authentication. In principle, this should allow easy switching of authentication and authorization service providers. Just as I can still keep the &#8220;label&#8221; for this site even if I move to a different web host, I can still keep my personal &#8220;label&#8221; no matter who I choose to manage my personal information.\u00c2\u00a0 So, now my universal username is <tt><tt>spiros.papadimitriou.name<\/tt><\/tt>, any emails sent to <tt><tt>spiros@papadimitriou.name<\/tt><\/tt> will find their way to me, you can call me on Skype using <tt><tt>spiros.papadimitriou.name\/call<\/tt><\/tt>, and so on.<\/p>\n<p>With such a unique identity tied to authorization and authentication services, the <a title=\"Giant Global Graph - Tim Berners Lee on DIG\" href=\"http:\/\/dig.csail.mit.edu\/breadcrumbs\/node\/215\">Giant Global Graph<\/a> and <a title=\"URLs are People, Too\" href=\"http:\/\/google-code-updates.blogspot.com\/2008\/02\/urls-are-people-too.html\">its materializations<\/a> would be one step closer to becoming really useful. If I want to use my identity to log and controll access to my data, I should be able to prove my claims.\u00c2\u00a0 Currently, FOAF and XFN allow assertion of relationshipt but provide no way to verify them.<\/p>\n<h4>Data portability<\/h4>\n<p>The point of this mental exercise so far is the following: <strong>A unique identity that can be verifiably associated with each and every data item that I produce is a prerequisite for making data ownership claims.<\/strong> Subsequently, we need to ask what fundamental rights should be associated with data ownership.\u00c2\u00a0 <strong>The first is the right to keep <em>my<\/em> information with <em>me<\/em><\/strong> or, in other words, &#8220;data portability&#8221;. Just as I can freely move my money from one financial institution to another, I should be able to move any of my information from one data warehouse to another.<\/p>\n<p>For example, consider my web search history. I don&#8217;t think I need to argue about the importance of historical information to improve search quality. If I decide for any reason to move to another search provider, I should be able to carry along all the information that&#8217;s directly associated with me.\u00c2\u00a0 This should include my search keyword history, as well as any <a title=\"Google SearchWiki brings custom search results - CNET\" href=\"http:\/\/news.cnet.com\/8301-17939_109-10102750-2.html\">additional information I may have contributed<\/a>.<\/p>\n<p>The actual details, however, may not be that straightforward.\u00c2\u00a0 Take, say, the third hit on a Google search.\u00c2\u00a0 Who is the &#8220;creator&#8221;?\u00c2\u00a0 Me by entering the search keywords, Google by producing the search results in response to those keywords, or the person who wrote the web page that contains them in the first place?\u00c2\u00a0 Similarly, when I buy gas, who is the &#8220;creator&#8221; of the transaction entry: me, Mobil, or American Express?<\/p>\n<p>Even though intuition can often be wrong, my intuitive response to the Google search example would be that both I and Google have an ownership claim on this particular search, which includes the query keywords as well as a ranking of URLs.\u00c2\u00a0 On the other hand, the person who wrote the contents of, say, the third URL has ownership claims only on those, and not the search results.\u00c2\u00a0 Furthermore, the thousands of people that provided feedback to Google&#8217;s ranking algorithms by clicking on this URL on similar searches have ownership claims on those searches, but not on mine.<\/p>\n<p>Finally, those two ownership claims (on keywords and on rankings) should probably not be treated the same.\u00c2\u00a0 If they were, then, say, MS Live could effectively copy Google by getting many users to move.\u00c2\u00a0 It seems reasonable to have the right to move my search history, but not the actual search results. However, I can imagine that some form of ownership claim on the rankings may be useful for other personal rights.<\/p>\n<p>This is a highly idealized example and I&#8217;m not sure what an appropriate litmus test for ownership is, but some form of legal consensus must be in place.<\/p>\n<h4>Transparency<\/h4>\n<p><strong>The second fundamental right is that I should know <em>who<\/em> is using my personal information and <em>how<\/em>. <\/strong>For example, if an insurance company accesses my credit history to give me a rate quote, I can find this out. It may not be a completely painless process but it is certainly possible today, with a regulatory framework that ensures this.\u00c2\u00a0 Similar regulations should be instituted to cover any and all forms of access to personal information.<\/p>\n<p>Data access should be fully transparent to all parties involved. If the an insurance company accesses my medical records, I should know this.\u00c2\u00a0 If the government does a background check on me, I should know this too.\u00c2\u00a0 Transparency is a prerequisite for accountability. Otherwise, individuals have very limited power to protect themselves from improper uses of their personal information.<\/p>\n<h4>Concluding remarks<\/h4>\n<p><strong>Much of the privacy research in computer science seems to assume that we can keep the existing legal and regulatory frameworks intact.<\/strong> Computer scientists taking such a position is even sadder than lawyers doing so; we have no excuse of failing to understand the technical issues.<strong> We cannot and should not make this assumption.<\/strong> Technical solutions should be subsidiary to new regulations.\u00c2\u00a0 But that doesn&#8217;t mean technologists cannot lead.\u00c2\u00a0 We should work towards supporting full transparency (for both individuals, as well as governments and corporations) rather than opacity and <strong>I&#8217;m currently in favor of a &#8220;shoot first, ask questions later&#8221; approach (and help lawmakers figure out the answers)<\/strong>. After all, if there is anything that the DRM wars have taught us, it&#8217;s that information really wants to be free. Why do we think it&#8217;s technically hard (to say the least) to prevent copying of music, movies and software but we still think it may be possible to prevent copying of personal information? As I pointed out in an <a href=\"http:\/\/www.bitquill.net\/blog\/?p=5\">older post<\/a>, it&#8217;s usually the use and not the possession of information that&#8217;s the problem.<\/p>\n<p>My point in this post is simple: <strong>we should not fight the wrong war<\/strong>. Instead, we need an easy way to make data ownership claims, and use this to enforce at least two fundamental rights: the ability to keep any personal data with us, and the ability to know who is using this data and how.<\/p>\n<p><strong>Postscript.<\/strong> This post was wallowing for a while as a draft (originally separated from <a href=\"http:\/\/www.bitquill.net\/blog\/?p=5\">this post<\/a>, then forgotten).\u00c2\u00a0 Since then, a recent <a title=\"Who Owns Your Friends? (MIT TR)\" href=\"http:\/\/www.technologyreview.com\/Infotech\/20920\/\">MIT TR article<\/a> discusses some aspects of data ownership.\u00c2\u00a0 Even better, I have since found an excellent <a title=\"Curating Yourself Online: What happens when your data is not your alone? (MIT TR)\" href=\"https:\/\/www.technologyreview.com\/Infotech\/20936\/\">short piece<\/a> in the same issue by <a title=\"Esther Dyson (Wikipedia)\" href=\"http:\/\/en.wikipedia.org\/wiki\/Esther_Dyson\">Esther Dyson<\/a>, with which I could not agree more.<\/p>\n<p><em><strong>Update.<\/strong> After posting this last night, I did some further Googling and found <a title=\"How Loss of Privacy May Mean Loss of Security - Scientific American\" href=\"http:\/\/www.sciam.com\/article.cfm?id=how-loss-of-privacy-may-mean-loss-of-security\">another piece by Esther Dyson<\/a> in the Scientific American. If you&#8217;ve read through my ramblings so far, then I&#8217;d urge you to read her article; she&#8217;s a much better writer than me, and has apparently been thinking about these issues for almost a decade, way before many people even knew what the Internet is. I should probably follow her more closely myself, as I agree disturbingly often with what I&#8217;ve read from her so far.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Every piece of content has a creator and owner (in this post, I will assume they are by default the same entity). I do not mean ownership in the traditional sense of, e.g., stashing a piece of paper in a drawer, but in the metaphysical sense that each artifact is forever associated with one or [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_monsterinsights_skip_tracking":false,"_monsterinsights_sitenote_active":false,"_monsterinsights_sitenote_note":"","_monsterinsights_sitenote_category":0,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":false,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[45],"tags":[21,18,15,19,58,16,17,23,20,4],"class_list":["post-7","post","type-post","status-publish","format-standard","hentry","category-scitech","tag-accountability","tag-authentication","tag-data-ownership","tag-data-portability","tag-opinion","tag-privacy","tag-provenance","tag-regulation","tag-transparency","tag-web"],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"","jetpack_shortlink":"https:\/\/wp.me\/p7x9xm-7","jetpack-related-posts":[],"jetpack_sharing_enabled":true,"_links":{"self":[{"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/posts\/7","targetHints":{"allow":["GET"]}}],"collection":[{"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/comments?post=7"}],"version-history":[{"count":1,"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/posts\/7\/revisions"}],"predecessor-version":[{"id":715,"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/posts\/7\/revisions\/715"}],"wp:attachment":[{"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/media?parent=7"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/categories?post=7"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/bitquill.net\/blog\/wp-json\/wp\/v2\/tags?post=7"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}