How Not to Build Algorithms

In the summer of 2016, President Duterte won the Philippine election by a landslide. The months before the election were grueling for many of us Filipinos, particularly on social media. Organized pro-Duterte efforts that skewed perceptions of candidates sprang up from seemingly nowhere, and online bots, trolls, and Facebook pages formed an army exploiting Facebook News Feed algorithms.

Websites containing fake or heavily biased news grew in number and were rampantly shared by Facebook pages with followings in the millions. Bots and trolls propagated posts and pages and threatened those who retaliated. Through this phenomenon, people were siloed into echo chambers devoid of productive conversation, and our efforts to fight online trolls and encourage conversation were futile by the end of the election. The plight and frustration of many Filipinos had gone too deep.

There were similarities to the US election as well, where News Feed algorithms reinforced cognitive biases and failed to filter out fake news. While Facebook claimed not to be a media channel, many of its users use it as such while unaware of the mechanisms behind it.

What are the News Feed algorithms like? Before 2012, the algorithm used to rank articles was called EdgeRank, which can be simplified into this equation:


where:

: edges
: relationship and proximity of user and content
: weight of user reaction
: decay parameter for time content was posted

More recently however, the number of features of their News Feed machine learning algorithm increased to the hundreds. Facebook now calculates a relevancy score, which is assigned to a post given a particular Facebook user. Then, on each user’s News Feed, the algorithm sorts Facebook posts by their relevancy scores. These scores are based on features such as interactions with other users, the user’s previous reactions to posts, and virality of the post - many of which limit a person’s Facebook experience to certain patterns rather than diversifying their experience.

It’s not just on the News Feed that algorithms have social or political (or other) implications. I recently came across a podcast of the Center for Democracy & Technology (CDT) on Cathy O’Neil’s research and book entitled, Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Here she discusses how models and algorithms can output biased decisions or be weaponized (humans create them after all).

In the podcast, she talks about the 2008 financial crisis. The mathematical models used in subprime lending were carelessly built (you could of course add greed into this equation). The models were created with faulty assumptions like, (1) residential real estate prices would always appreciate and (2) mortgage defaults were independent and identically distributed. Eventually what resulted was the biggest financial crisis since the Great Depression.

She also described how policing algorithms have wrongfully labeled people as high risk offenders. Arrests are used to predict future crimes committed, and police are usually deployed to areas where more arrests took place. But arrests don’t necessarily equate to crime, and so this whole process contributes to uneven policing.

Algorithms that calculate recidivism rates also have inherent biases in them. In a number of criminal courts in the United States, black Americans were twice as likely to be wrongfully labelled as high risk offenders as white Americans. This came as a result of biased training data and the ensuing trade-off between predictive accuracy and level of discrimination.

Facebook’s “Ethnic Affinity” targeting option is also controversial. Facebook algorithms predict user ethnic affinity without their consent, and companies are able to use this for targeted marketing. Corinthian Colleges, a U.S. for-profit college, used this method and aggressively targeted a number of low-income black single mothers and other vulnerable groups and convinced them that a degree was their shot at a better life. Instead, these people were left entrenched in student debt.

These examples are not the only cases where algorithms had or might have negative consequences. Discrimination in the workplace, experiments without user consent, among others happen(ed) as well.

We have an enormous responsibility as future data scientists. We’ll be creating and wielding algorithms and models that will affect people’s lives, so we must be mindful and find good guidelines so that people are protected against harmful decisions and labeling. We should try to avoid pitfalls and keep others’ models in check because we might just be heard (after CDT and other civil rights organizations expressed their concerns, Facebook amended their ethnic affinity policy).

We live in a time where data science is widely used and glorified. Companies, people, and institutions increasingly rely on models and algorithms to make even the most important decisions. However, left unchecked, these models and algorithms can inadvertently threaten democracy, increase inequality, and negatively impact people’s lives. We data scientists must try our best to prevent that from happening.