Social Media Platforms Are Training Their AI Models on Your Content, but You Can Stop Them (Sometimes)
Social Media Platforms Are Training Their AI Models on Your Content
Recently, the Amazon-owned Twitch made headlines for announcing that it would allow users to opt out of having their streams, VODs, and chats fed into an AI training machine. This, understandably, made a lot of people upset. But, it turns out, Twitch might be one of the more moderate social media sites on this front.
Most social media apps are training some kind of AI on your data, and few make it as easy as a single toggle to opt out at all. Almost every social media site out there makes it extremely hard to know exactly how your data is used these days.
Training generative AI models like Google’s Gemini are often lumped in with more mundane (but still machine learning-powered) features like, say, YouTube’s algorithm. Sifting through privacy policies to even find out whether a social app contributes to the kind of generative AI that’s proven so controversial is an undertaking that would deter most lawyers.
Facebook’s AI Training Policy
Facebook’s parent company Meta is pivoting to AI and rushing to catch up to companies like OpenAI, Anthropic, and Google in the process. The company’s policy on training AI with your data is extremely broad, not only encompassing your posts, photos, and interactions on Facebook, but also data collected from third-party brokers, and “information that is available on the internet.”
That’s a phrase so all-encompassing that it’s hard to imagine any data Facebook is technically capable of scooping up that it would refrain from collecting. Facebook says it stops just short of training AI on private messages with friends or family, “unless you or someone in the chat chooses to share those messages with our AIs.”
Unfortunately, unless you’re in the EU, Facebook makes it impossible to opt out of training AI with your data. The company carves out a narrow exception (where it is obligated to by law) for the specific scenario where you find personally identifying information about you included in a response from one of its AI tools.
However, keep in mind that Meta considers interacting with Facebook’s AI tools in the first place as consent to train future AI models on your interactions with it. So, even trying to figure out whether your personal information has been scooped up at all could expose you further.
Other Social Media Platforms’ AI Training Policies
Instagram is owned by Facebook’s parent company Meta, so many of its policies are the same as what we covered for Facebook in the section above. However, Instagram has some specific issues of its own, such as the Muse feature, which briefly let users create AI-generated images of other users without their consent, before immediately removing that feature after realizing it was a terrible, horrible idea.
LinkedIn is owned by Microsoft, and Microsoft famously is all-in on AI. According to LinkedIn’s official policy, users’ posts, comments, profile data, resumes, and group activity, among other types of data, can all be used to train AI models. Unofficially, it should be noted that last year LinkedIn was sued over allegations that it was training AI on private DMs – an allegation LinkedIn denies.
Reddit is complicated: The company doesn’t train its own AI models, but it does have deals with both OpenAI and Google to train their models on its data. Reportedly, Reddit has at least considered ending those deals (at least in part because, like everyone else, the company has seen a drop in traffic as AI eats search’s lunch).
However, what’s clear is that most social media platforms are using your data to train their AI models, and it’s often not easy to know exactly how your data is being used or to opt out of it. It’s a complex issue that requires a closer look at each platform’s policies and practices.