Winter term 2026 NPFL112

YOUR MAIN RESOURCES

Lecture 1.

Date: 2026-10-02

Goals

  • access the Jupyter Hub RStudio (save your credentials)
  • access DataCamp (create your account with the same e-mail address with which you enrolled)
  • Overview of
    • the course scope
    • grading requirements
  • Know where to find presentations online:
    • https://ufal.github.io/NPFL112/
    • Press s to display speaker view with ample fluent notes
  • Know how to download presentations in other formats (pdf, html pages) from RStudio on Jupyter Hub
    • SCRIPTS.NPFL112/slides/slides_html
    • SCRIPTS.NPFL112/slides/slides_pdf
Starting Positions Mentimeter

Participants’ link: https://www.menti.com/alxeutazt9en

Admin needs to log in at Mentimeter (and open it at MyMentis > Dashboard > NPFL112_PersonalLearningGoals > … > Share with Participants) here:

https://www.mentimeter.com/app/presentation/alz5fpjhnz8kjfbe447psbvptz9pupr6/edit

Presentations

https://ufal.github.io/NPFL112/slides/01_Introduction.html static html pdf

https://ufal.github.io/NPFL112/slides/02_HowToRStudio.html static html pdf

Activities

In R Studio

  1. Log in at RStudio

  2. In the Files tab (right bottom pane), create a new directory (folder) and call it - exactly! - SUBMISSIONS.NPFL112 .

  3. Explore the folders SCRIPTS.NPFL112 , DATA.NPFL112

  4. Local download example: Download to your computer ~/SCRIPTS.NPFL112/slides/slides_pdf/04_NavigatingRStudioForProgramming.pdf.

  5. Log in at https://www.datacamp.com/users/sign_in (you should have received an invite from the DataCamp system on Tuesday Sep 29 or later). Sign up with the same e-mail address with which you have enrolled in this course. If there is time left, start with your first home assignment.

Assignments

Assignment 1: Deadline: Fri Oct 9 , 9:30 AM

In RStudio, make sure that you have created your folder called SUBMISSIONS.NPFL112.

Open the folder HOMEWORK_ASSIGNMENTS. Select (tick) the file HW_001.R and copy it into your SUBMISSIONS.NPFL112 folder.

If you were present in the first lecture and received a printout of the file, write your responses directly into the printout. Please use a legible handwriting!

Open this file, read the instructions and proceed accordingly. Go through the entire file but do not spend more than 30 minutes with it. You are going to exchange the hard copy of your solution with a peer in the next session (hence the legibility appeal). If you are more comfortable typing straight in the file rather than writing in the paper sheets, please print out the result with no more than one page per A4 page (keep the font size eye-friendly). Be ready to explain for your peer what you know (or guess) about the code.

Assignment 2: Deadline: Fri, Oct 9, 9:30 AM

Lecture 2.

Date: 2026-10-09

Goals

  1. Internalize the following concepts:

    1. Data types/classes (numeric, character, boolean)

    2. Data type coercion (what happens to your numeric vector when you blend in a non-digit, etc. )

    3. Data structures (vectors, data frames/tibbles)

    4. Functions and their arguments

    5. Working directory

  2. Learn to invoke and read the built-in R Help

  3. Open a Quarto (.qmd) file and run code chunks manually

Presentations

https://ufal.github.io/NPFL112/04_NavigatingRStudioForProgramming.html static html pdf

https://ufal.github.io/NPFL112/05_VariablesFunctions.html static html pdf

https://ufal.github.io/NPFL112/06_WorkingDirectory.html static html pdf

Activities

  1. Lecture start: group work, exchange about HW_001.R: During your homework, which structures/patterns caught your eye?
  2. Hands-on together: Copy the file SCRIPTS.NPFL112/05_VariablesFunctions.qmd to your home directory.
    1. Open it and run (execute) all code chunks. Watch what happens.
    2. When you are done with executing the chunks, find the “Render” button. Explore the rendering options. Render the file as an html file. It will appear in the same directory where you had the corresponding qmd file and will inherit its name. Open the resulting html file in a web browser.
    3. Open a new Quarto .qmd file. Render it by hitting the Render button.

Assignments


Lecture 3.

Date: 2026-10-16

Goals

  1. Internalize the following concepts:
    1. Vector recycling
    2. Vector element vs. vector position index
    3. logical operators (>,<, >=, <=, !=, ==, &, |)
  2. Know how to:
    1. Extract elements from a vector (by position or by a condition expressed by a logical operator).
    2. Extract values from data frames by rows and columns in base R.
    3. Read a tabular file (.csv) into a data frame object.
    4. Write a data frame object into a tabular file (.csv)
  3. Recognize transformations of tabular data (no coding): filtering rows, selecting columns, aggregations, aggregations in groups

Presentations

Finish https://ufal.github.io/NPFL112/05_VariablesFunctions.html (from slide No Coercion with errors).

Activity to new topic

table transformations, see below

https://ufal.github.io/NPFL112/07_Exploring_dataframes.html static html pdf

Activities

Look at the (printed) tables and keep them ready at your hands. The tables are bits of gapminder data, also file clips_handouts.pdf in SCRIPTS.NPFL112 and https://ufal.github.io/NPFL112/clips_handouts.pdf. Each sheet contains a pair of tables. What do you need to do to the table on the left to obtain the table you see to the right? Try your luck in a multiple choice test here: https://quest.ms.mff.cuni.cz/class-quiz/quiz/NPFL112_02_01_clips

and check out whether you can describe most of the transformations - in your own words, no scientific jargon, no coding. Note the difficult ones and ask about them after the exercise. Duration: 10 - 15 minutes.

Assignments

DH-Certificate Assignment No 1, deadline Thu Oct 22, 16:00

In RStudio, produce a markdown file containing a list of paths leading to the folders containing the datasets available for our work. Render it as html.

Assignment 4: Deadline: Fri Mar 20

https://app.datacamp.com/learn/courses/introduction-to-the-tidyverse?cf-exp=comm--ai-native-cdp-layout%3Aon


Lecture 4

Date: 2026-10-23

Goals

Presentations

https://ufal.github.io/NPFL112/08_DiversePlots.html static html pdf

Activities

Mentimeter WarmUp: https://www.menti.com/

Teacher’s GUI: https://www.mentimeter.com, presentation Dplyr and ggplot2 for beginners. Needs login and the presentation must be launched.

Your first data report in Quarto: tracking the world’s billionaires

  1. Open a new Quarto (.qmd) file (a Document, but Presentation would be fine, too).

  2. In the dialog window, give the document the title “Billionaires Investigation” and fill out your name.

  3. Save the file into SUBMISSIONS.NPFL112 under the name my_first_data_report.

  4. Check out the YAML header. Add today’s date to it with this line: date: today

  5. Hit the Render button (top of the pane). If you did it correctly, RStudio’s bottom right pane expands with the Viewer tab fronted, showing your rendered html file. Switch to the Files tab and find the resulting html file that Viewer is showing you. Can you see the date, too?

  6. Get back in the top-left pane and start editing your Quarto document. Have a look at the template in both the Source and the Visual editor mode (toggle these options in the top left corner). Particularly inspect the Running Code section and hit the green arrow/triangle to run the code inside. Once you have read the text in the template, you can erase it.

  7. Insert a code chunk with Ctrl + Alt + i or with the green +C button in the top right corner of your pane.

  8. Type into the chunk the code to mount/activate the necessary libraries dplyr , readr, and ggplot2.

    Call all tidyverse libraries simultaneously

    You can as well load all tidyverse libraries simultaneously by calling library(tidyverse).

  9. In the same chunk, create a variable called df (to remember that it is a “data frame”) by reading in the file with the following path: ~/DATA.NPFL112/billionnaires_combined.tsv . It is a .tsv file (tab-separate values), so you need the read_tsv function from the readr library.

    Understand the file-reading functions of the readr library

    There are dedicated functions for specific formats, such as the U.S. .csv, the European .csv, and .tsv. All these formats capture tables in plain text by a convention of separating columns with a specific delimiter character (comma, semicolon, or tabulator). They all are derived from the read_delim function. You can try it out with this very file. Mind to set the delimiter argument to tabulator. Check out the other arguments of these functions in their documentation!

  10. Explore the data frame df. Create another code chunk and print first ten rows of this data frame.

  11. Overwrite df so that it will only contain data from the year 2022.

  12. Create yet another code chunk. In this code chunk, write the code to print a histogram of the wealth distribution among the billionaires (column daily_income). Experiment with the binwidth argument of geom_histogram.

  13. Now create a new code chunk that prints a scatterplot with age on the x-axis and daily_income on the y-axis. Can you observe any relation between age and daily income? E.g. do senior billionaires appear to earn more per day than the younger ones?

  14. Above each chunk with a plot add a heading formatted as Heading 2 and a short text describing for your audience what the given plot demonstrates.

  15. Render your document without changing anything in the YAML header.

  16. Now replace html with pdf in the YAML header and hit the Render button again. If you get an error message, install a package called tinytex and try again without explicitly calling/mounting/activating it. RStudio ought to do it by itself. If you do not get a pdf file, give up. Make a note about this failure in your file and re-render it as html. That always works.

Assignments

Assignment 5: Deadline: Mar 27

Finish the billionaires exercise from this lecture and place it into your SUBMISSIONS.NPFL112 folder.

Assignment 6: Deadline: April 8, noon

Do this exercise and place the resulting file (R or .qmd from a Quarto file) into your SUBMISSIONS.NPFL112 folder. If you produce a .qmd file (you are encouraged to!), please also add a rendered file (html or pdf).

Lecture 5

Date: 2026-10-30

Goals

Wrap up operations on a single table with dplyr.

Presentations

https://ufal.github.io/NPFL112/09_Aggregations_with_dplyr.html static html pdf

https://ufal.github.io/NPFL112/11_Computations_mutate_with_dplyr.html static html pdf

Activities

Look again at the worksheets with data frames. Do not take them out of their plastic sleeves. From the label sheet, pick for each data frame sheet a label that encodes the transformation of the left table into the right table.

Assignments

Lecture 6

Date: 2026-11-06

  • train reading a foreign code - recall and apply the bits familiar to you, infer the meaning of new elements from the context

Activities

Group work on paper: interpret someone elses’ code in greatest possible detail. Hazard guesses! The code has one issue in it. Will you find it? Compare your results with other groups. In-person lecture: code in hardcopy. Here is a link to the code in pdf.

Here is the clue to the exercise - the code broken down into meaningful chunks and explained. static html pdf

Lecture 7

Date: 2026-11-13

Deadline for HW02 was April 8, noon.

Goals

  • Internalize the concept of relational database

    • primary key, secondary key

    • one-to-one, one-to-many…

    • item pointing to itself

  • dplyr : join dataframes on a pair of columns

    • inner join vs. left/right join vs. full join

    • semi join, anti join

Presentations

https://ufal.github.io/NPFL112/12_JoiningDplyr.html static html pdf

Activities

Assignments

Lecture 8

Date: 2026-11-20

Goals

  • string search with stringr - just enough to understand why it is relevant when restructuring data frames with tidyr

    • identify just the first match, all matches

    • detect/count/match/extract/replace

    • regular expressions: any character, interval, groups (a few random features, you will have to learn regular expressions separately)

  • get to know the tidyr library

  • tidy data - longer vs. wider format

    • pivot_longer, pivot_wider
  • separate or unite columns based on substrings in their values

  • separate rows

Presentations

https://ufal.github.io/NPFL112/15_tidyr_stringr.html static html pdf

Activities

Self-paced exploration of operations on strings with stringr on Mentimeter https://www.menti.com/alwuzuf2md8h (expires Apr 18)

Assignments

Lecture 9

Date: 2026-11-27

Goals

  • Parse a JSON file to a list or a data frame

  • Extract information from nested data frame columns

Presentations

https://ufal.github.io/NPFL112/17_from_json.html static html pdf

Activities

You can try and download this file: https://api.fbi.gov/\@wanted into a folder where you have writing permissions. Call the file whatever you want. The correct suffix is .json.

my_destination_file <- "~/DATA.NPFL112/2026-FBI-wanted.json"
download.file(url = "https://api.fbi.gov/@wanted", 
  destfile = my_destination_file, 
  mode = "wb") 
#Use wb with Excel files on Windows, here not necessary
library(jsonlite)
library(tidyverse)
fbi_simple_list <- jsonlite::fromJSON(my_destination_file, 
  simplifyDataFrame = TRUE, 
  simplifyMatrix = TRUE, 
  simplifyVector = TRUE, 
  flatten = TRUE)
# with complex structures, explore calling arguments of str() to display just a little
str(fbi_simple_list, max.level = 2, vec.len = 0, list.len = 10)
List of 3
 $ total: int NULL ...
 $ page : int NULL ...
 $ items:'data.frame':  20 obs. of  43 variables:
  ..$ pathId                : chr [1:20]  ...
  ..$ uid                   : chr [1:20]  ...
  ..$ title                 : chr [1:20]  ...
  ..$ description           : chr [1:20]  ...
  ..$ images                :List of 20
  ..$ files                 :List of 20
  ..$ warning_message       : chr [1:20]  ...
  ..$ remarks               : chr [1:20]  ...
  ..$ details               : chr [1:20]  ...
  ..$ additional_information: chr [1:20]  ...
  .. [list output truncated]

How many wanted persons are there in the items data frame? You can tell from the str output above.

fbi_complex_list <- jsonlite::fromJSON(my_destination_file, 
  simplifyVector = FALSE, 
  simplifyMatrix = FALSE, 
  simplifyDataFrame = FALSE,
  flatten = FALSE )
str(fbi_complex_list, max.level = 3, vec.len = 1, list.len = 4)
List of 3
 $ total: int 1257
 $ page : int 1
 $ items:List of 20
  ..$ :List of 43
  .. ..$ pathId                : chr "https://api.fbi.gov/@wanted-person/c034fc96d059474f97acdbeed3dd20c9"
  .. ..$ uid                   : chr "c034fc96d059474f97acdbeed3dd20c9"
  .. ..$ title                 : chr "MARY VIRGINIA GARCIA WEBB - BILLINGS, MONTANA"
  .. ..$ description           : chr "June 8, 1996\r\nBillings, Montana"
  .. .. [list output truncated]
  ..$ :List of 43
  .. ..$ pathId                : chr "https://api.fbi.gov/@wanted-person/98df11f3903d4ecab8940693471cbca8"
  .. ..$ uid                   : chr "98df11f3903d4ecab8940693471cbca8"
  .. ..$ title                 : chr "CARLOTTA MARIA SANCHEZ"
  .. ..$ description           : chr "August 30, 1979\r\nTaholah, Washington"
  .. .. [list output truncated]
  ..$ :List of 43
  .. ..$ pathId                : chr "https://api.fbi.gov/@wanted-person/fca89f1133a74259b4a0f944cd238343"
  .. ..$ uid                   : chr "fca89f1133a74259b4a0f944cd238343"
  .. ..$ title                 : chr "ELSIE ELDORA LUSCIER"
  .. ..$ description           : chr "August 30, 1979\r\nTaholah, Washington"
  .. .. [list output truncated]
  ..$ :List of 43
  .. ..$ pathId                : chr "https://api.fbi.gov/@wanted-person/461a97c6a7214afeb22c8b70d83a5fee"
  .. ..$ uid                   : chr "461a97c6a7214afeb22c8b70d83a5fee"
  .. ..$ title                 : chr "SEAN E. BROWN"
  .. ..$ description           : chr "Murder/Intent to Kill/Injure"
  .. .. [list output truncated]
  .. [list output truncated]

Assignments

Assignment : 9 Deadline: May 14

Matrix, List (Intro to R)

Condition, Loop, Function

Iteration over a list with purrr

  • alternatively, you can learn the apply functions from base R, but this is more comfortable: possibly easier to learn and (unlike apply) would never surprise you with transforming your data frame to a matrix. The entire purrr course is worth working through!

https://app.datacamp.com/learn/courses/foundations-of-functional-programming-with-purrr