The old version of this page said I had studied six sets of Filipino text and built models that got 87 out of 100 right.
I had not. So I opened one set properly instead, and it turned out to be more interesting than six I never touched.
First: You Will Not See Any Of It
The set is 27,383 posts from our 2016 and 2022 election seasons. Each one is marked "hate" or "not hate".
Almost every write-up like this shows you a few examples so you can see what the words look like. I am not going to.
These posts are real abuse aimed at real, named people. Putting them here to make a chart look interesting just spreads them further, in nicer type.
So there are no examples. There is also no "most common words" list, because on this set that list is mostly slurs.
The only words I counted are little joining words. Ang. The. Sa. Of. Words that carry no meaning on their own. Everything I found still works without showing you a single post.
The Test Is Leaky
Sets like this come split into parts. You teach a computer with one part, then test it with another part it has never seen. That is the whole point. It is like giving a student a test with different questions than the practice sheet.
I checked whether the parts really are different. They are not, quite.
273 posts show up in more than one part.
So some of the test questions were on the practice sheet. Any score anyone reports on this set is a little too high, and nobody knows by how much.
I did not quietly delete those posts. If I clean it up here, my set stops matching the one everybody else uses, and the problem just hides. So I counted it and wrote it down.
A Ghost In The Shape
I measured how long the posts are. The chart has a cliff in it. Lots of posts up to 140 letters, then it falls off.
79.03 out of every 100 posts fit inside 140 letters.
That is not about how Filipinos write. For years Twitter only let you type 140 letters. In late 2017 they doubled it.
This set covers 2016 and 2022 — one election on each side of that change. So the shape of my chart is really a picture of a company's rule, not of a country's writing.
I like this one. If I had not known that bit of history I would have written something silly about Filipinos preferring short posts.
The Finding I Am Least Sure Of
Here is the strongest pattern, and the one I want you to be careful with.
The posts marked "hate" use far more Tagalog. Of their joining words, 64.79 out of 100 are Tagalog. For the other posts it is 49.32.
And 47.01 out of every 100 hate posts show no English joining words at all. For the rest it is 27.03.
That is a big gap. Now, what does it mean?
It could mean people insult each other in Tagalog and stay polite in English. That sounds believable.
It could also mean the people who marked these posts were quicker to call a Tagalog post hateful. That sounds believable too.
My data cannot tell these apart. Not "I have not checked yet" — it genuinely cannot, because I only have the words and the label, not the person who wrote either.
And this matters. If you teach a computer on this set, it learns that pattern either way. Including if the pattern is really about the markers, not the writers.
How I Measured Language
My method here is crude and I want to be upfront about it.
I wrote two lists of joining words, one Tagalog and one English, before I looked at the data. Then I checked which lists each post matched.
It misses a lot. Borrowed words, misspellings, and the way we bend English words with Tagalog endings. So when I say at least 25.99 out of 100 posts mix both languages, treat that as the floor, not the answer.
11.54 out of 100 posts matched neither list. Short ones, hashtags, names, heavy slang. I left that number visible instead of quietly spreading it into the others to make the pie look tidy.
Want to see all the charts and data tables?
View the Full Analysis →