2.5 Comparing the three sentiment dictionaries
Lexicons have diff quality
Use three sentiment lexicons and examine how the sentiment changes across the narrative arc of Pride and Prejudice.
<- tidy_books %>%
pride_prejudice filter(book == "Pride & Prejudice")
pride_prejudice
## # A tibble: 122,204 × 4
## book linenumber chapter word
## <fct> <int> <int> <chr>
## 1 Pride & Prejudice 1 0 pride
## 2 Pride & Prejudice 1 0 and
## 3 Pride & Prejudice 1 0 prejudice
## 4 Pride & Prejudice 3 0 by
## 5 Pride & Prejudice 3 0 jane
## 6 Pride & Prejudice 3 0 austen
## 7 Pride & Prejudice 7 1 chapter
## 8 Pride & Prejudice 7 1 1
## 9 Pride & Prejudice 10 1 it
## 10 Pride & Prejudice 10 1 is
## # … with 122,194 more rows
<- pride_prejudice %>%
afinn inner_join(get_sentiments("afinn")) %>%
group_by(index = linenumber %/% 80) %>%
summarise(sentiment = sum(value)) %>%
mutate(method = "AFINN")
<- pride_prejudice %>%
bin_pride_prejudice inner_join(get_sentiments("bing")) %>%
mutate(method = "Bing et al.") %>%
count(method, index = linenumber %/% 80, sentiment) %>%
pivot_wider(names_from = sentiment,
values_from = n,
values_fill = 0) %>%
mutate(sentiment = positive - negative)
## Joining, by = "word"
bind_rows(afinn,
%>%
bin_pride_prejudice) ggplot(aes(index, sentiment, fill = method)) +
geom_col(show.legend = FALSE) +
facet_wrap(~method, ncol = 1, scales = "free_y")
- The lexicons for calculating sentiment give results that are different in an absolute sense but have similar relative trajectories through the nove