Commentary

Privacy, With an Asterisk

Earlier this May, Darren Zhou was arrested for writing threats to kill his ex-girlfriend. Zhou shared detailed plans with ChatGPT, which OpenAI, the creator of the popular large language model, noticed and eventually brought to the attention of legal authorities. 

Though the fact that Artificial Intelligence (AI) may now play significant roles in criminal cases may seem fascinating to the public, it is much less important than the question of privacy that this case entails. The fact that Zhou was flagged means that AI companies may be monitoring chats. In fact, OpenAI stated that data may be disclosed “to protect the safety, security, and integrity of our products, employees, users, or the public.” However, as people gradually begin to incorporate AI models and tools in their daily lives, Zhou’s case raises the question of where the boundary between violation of personal privacy and necessary intervention lies.

There are theories that staff from companies like OpenAI are secretly watching all AI conversations and documenting every single word choice. However, most major AI companies do not manually monitor chats, instead implementing an automatic detector for harmful conversations or activities, such as in the case of Zhou. Each company may have a slightly different detection policy, but most pay attention to illegal or dangerous activity. After an alert, a human staff reviews the report to determine whether it was a false positive or a real threat. In the process of doing so, some companies like OpenAI can involve legal authorities.

Therefore, it is also a common misconception that conversations with AI models remain completely internal between the model and the user. However, when the user sends a message to a model like ChatGPT, the message is stored on cloud servers managed by the company operating the model. In the case of companies like OpenAI, the messages may then be used to maintain chat history, improve the model, or check for policy violations.

The issue of AI privacy is surprisingly important to students, especially teenagers. Research from Pew Research Center showed that almost 1 in 8 students used AI for emotional support. If privacy in AI usage for emotional support, which teens likely resort to due to the lack of a confidant to rely on, isn’t assured, those suffering emotional struggles would no longer have the only place they could comfortably rely on. It is still painfully true, however, that setting blind exceptions for cases like these can also create exploitable loopholes in policy. The answer is therefore not blind leniency nor blind strictness.

In order to answer a question about privacy, there must first be a definition of privacy, or at least what it requires. Oftentimes, the privacy of information is determined by who has access to the information and who it can be shared with. However, what often gets neglected is considering whether the owner of the information is aware of how much of their privacy is protected. Since there aren’t any universal laws in privacy, along with other fields of ethics, settling on one specific objective set of values is most likely impossible. Therefore, decisions about who has access to personal information will ultimately rest with the model and the company, as they develop policies that they believe strike the optimum balance. However, users of such services should be informed about how much individual information can be shared with whom. 

To allow such a transparent environment to exist, users must make an effort. The most accessible and actionable method is actually reading the lengthy privacy policy. Yes, it’s extremely boring to read through twenty pages of lengthy legal statements, but even if we don’t read its entirety, skimming through the parts we deem important—including where and how our data is used—should be worth it to prevent feeling deceived or having our privacy violated in the future. 

Corporations also have major responsibilities. Companies, even outside of the context of AI, often use vague, general language, such as “protect against harm to the rights, property or safety of Google, our users or the public” instead of specifying the exact type of harm with examples like “threats of theft, murder, or violent illegal behavior.” It is convenient for corporations to use such language in order to escape as much legal responsibility as possible, but such wording does not add any useful information. If a company cannot explain its intentions, how is it reasonable to assume that customers can understand them? Though it’s sometimes necessary to use legal jargon, it should be the company’s duty to optimize transparency to protect its users. 

If such foundations—about how user data is used and shared—are clear, the balance may adjust towards a more optimal balance, wherever that may be. On one hand, a company that overfocuses on safety and neglects privacy will struggle to attract users, forcing it to either adjust its policy or risk falling behind in the market. On the other hand, a company obsessed with concealing all user activity, legal or otherwise, at the expense of safety may face numerous lawsuits for allowing such dangers to reach the public. Telegram is a good example of such harm-inducing platforms. Though not yet experiencing a decline in usage, Telegram, with its heavy emphasis on user privacy, has repeatedly clashed with governments in costly legal disputes. While Telegram’s decentralized structure may allow it to withstand such pressure, an AI company with a clearly identifiable headquarters and a centralized database would likely be far more vulnerable to government intervention. If AI corporations don’t work towards finding this balance, they will inevitably face consequences. In a nutshell, transparency gives the user the ability to judge whether they are willing to accept the balance their service has set. Giving users such abilities may create positive side effects, such as market pressure that rewards companies for working towards a favorable balance while discouraging those that sway toward either extreme.