CS U111 · Lecture 9 · Notes: the revision map
Strings, in one page
Rule cards, three worked examples at exam level, the slide corrections, and the trap table. First time with strings? Start with the lesson, which has the input widget.
Scope: Lecture 9, Strings in C. Reading: Hanly & Koffman ch. 8. A string is a char array plus one rule: it ends at a 0 byte, '\0'. Everything from Lecture 8 (indexing, traversal, swapping, "C never checks bounds") carries over unchanged. So the new material is that rule, three ways to read input, and about eight library functions. It's new to everyone in the batch: school Python and C++ hide all of it.
Exam relevance. The past mid-sem papers already ask char-array questions that walk to '\0': 2024 Part 2 Q8 (count matching characters of two arrays) and 2024 Part 2 Q11 (read characters until Enter, a 10-mark program). If Lecture 9 comes before this year's mid-sem, expect it to be examinable. Not covered yet: passing strings to functions, strtok (both after pointers), and the deeper library tour (Lecture 10).
The whole thing in one idea
A string is a char array whose text ends at the first 0 byte. "Hi" in quotes is 3 bytes: 72, 105, 0. Every string function (and %s) walks from box 0 until it meets that 0. The array's size never enters into it. So the size must be at least length + 1, and moving the 0 changes the string.
| Declaration | What you get (compiled and printed byte by byte) |
|---|---|
char g[6] = "Hi"; | 72 105 0 0 0 0: unused boxes are 0, not garbage |
char g[] = "Hi"; | Size counted for you: 3 (sizeof = 3, strlen = 2) |
char name[20]; (inside a function) | 20 unknown bytes, possibly with no 0 at all. Not "empty": don't print it until you've stored something |
char g[2] = "Hi"; | No room for the 0, so not a string. clang: "initializer-string for character array is too long" |
name = "Ada"; (after declaring) | Compile error: "array type 'char[10]' is not assignable". Use strcpy(name, "Ada"); |
Rule cards
1 · Walk to the 0, not to a size
Every hand-written string loop.
for (int i = 0; s[i] != '\0'; i++) { … s[i] … }
Counting the passes of this loop is strlen. Avoid i <= strlen(s): one pass too many.
2 · Reading one word: scanf("%Ns", s) with N = size − 1
Names without spaces, codes, single words.
char name[10];
scanf("%9s", name); /* no & for an array; skips leading spaces; stops at a space */
Without the 9, a long word overflows the array: undefined behaviour.
3 · Reading a whole line: " %N[^\n]" or fgets
Anything that may contain spaces.
scanf(" %29[^\n]", line); /* leading space skips a leftover '\n' */
fgets(line, sizeof(line), stdin); /* bounded, but KEEPS the '\n' */
Scansets in general: %[^,] stops at a comma, %[0-9] takes only digits. gets: never (removed from C in 2011).
4 · Strip the newline fgets keeps
After every fgets whose result you compare or print mid-line.
for (int i = 0; line[i] != '\0'; i++)
if (line[i] == '\n') { line[i] = '\0'; break; }
Tested: "Ada\n" becomes "Ada", and strlen goes from 4 to 3.
5 · The string.h four
#include <string.h>
strlen(s) /* characters before the 0 (unsigned: print with (int) cast) */
strcpy(dest, src) /* copy INCLUDING the 0: the way to "assign" a string */
strcat(dest, src) /* append src at dest's 0 */
strcmp(a, b) /* 0 if equal; <0 if a sorts first; >0 if a sorts after */
None of them checks dest's size. Before copying or joining, check that strlen(dest) + strlen(src) + 1 is at most the array's size.
6 · Equal text means strcmp(a, b) == 0
Every string comparison.
if (strcmp(a, b) == 0) printf("same\n");
a == b compares addresses: always false for two arrays. if (strcmp(a, b)) on its own is true when they differ.
7 · One character at a time: ctype.h
#include <ctype.h>. You supply the loop.
isalpha(c) isdigit(c) toupper(c) tolower(c)
s[i] = toupper(s[i]); /* in place; non-letters unchanged */
The is… functions return nonzero for true, not necessarily 1: test if (isdigit(c)).
8 · Pair the ends: reverse and palindrome
Anything symmetric. It's Lecture 8's reverse-an-array.
for (int i = 0; i < len / 2; i++) /* partner of i is len - 1 - i */
Stop at len / 2. Swapping all the way to len swaps every pair twice and undoes the reversal.
9 · Count starts, not spaces: the in-word flag
Words, runs, groups: anything separated by one or more delimiters.
if (s[i] != ' ' && !inWord) { count++; inWord = 1; }
else if (s[i] == ' ') inWord = 0;
Three worked examples at evaluation level
Worked example 1
The in-class activity: palindrome checker
Given char word[20] = "level";, decide whether it reads the same forwards and backwards. Test on "level" and "lovelace".
- Name the pattern. Card 8: pair box
iwith boxlen − 1 − i. And a flag (Lab 4): assume it's a palindrome, and try to disprove it. - Get the length.
int len = strlen(word);. Use the text length (5), not the array size (20). With 20, the pairs would compare letters against the zero-filled tail. - Write it.
#include <stdio.h> #include <string.h> int main(void) { char word[20] = "level"; int len = strlen(word); int isPal = 1; /* assume yes, try to disprove */ for (int i = 0; i < len / 2; i++) { if (word[i] != word[len - 1 - i]) { isPal = 0; break; } } if (isPal) printf("%s is a palindrome\n", word); else printf("%s is not a palindrome\n", word); return 0; } - Trace "level" (len 5, so
len / 2= 2 anditakes the values 0 and 1):Output:i pair compare result 0 (0, 4) lvslmatch 1 (1, 3) evsematch — loop ends; the middle v(box 2) has no partner and needs none.isPalis still 1.level is a palindrome. - Trace "lovelace" (len 8, so
iwould run 0–3):i = 0comparesl(box 0) withe(box 7). That's a mismatch, soisPal = 0andbreak, after one comparison. Output:lovelace is not a palindrome. - Edge cases an examiner likes. Even length has no middle: "abba" checks (0, 3) and (1, 2), and prints palindrome. Case matters: "Level" is not a palindrome to this code, because
'L'(76) ≠'l'(108). To ignore case, comparetolower(word[i]) != tolower(word[len - 1 - i])(tested: "Level" then gives 1). State which you assume.
All four outputs (level, lovelace, Level, abba) are from the compiled program.
Worked example 2
Word count with the in-word flag, on messy spacing
Count the words in char s[] = " Ada wrote code ";. It has two leading spaces, double and triple gaps, and a trailing space.
- Why the obvious idea fails. "Count the spaces and add 1" gives 8 spaces + 1 = 9. The answer is 3. Spaces aren't words, and runs of spaces confuse any count based on them.
- Count word starts instead (card 9). A start is a non-space character that comes when
inWordis 0.int wordCount = 0, inWord = 0; for (int i = 0; s[i] != '\0'; i++) { if (s[i] != ' ' && !inWord) { wordCount++; inWord = 1; } else if (s[i] == ' ') { inWord = 0; } } printf("%d\n", wordCount); - Trace (repeated rows grouped):
i s[i] what happens inWord count 0–1 ␣ ␣ space, already outside 0 0 2 A word starts: count it 1 1 3–4 d a inside a word: neither branch runs 1 1 5–6 ␣ ␣ space: leave the word (the second space changes nothing) 0 1 7 w word starts 1 2 8–11 r o t e inside 1 2 12–14 ␣ ␣ ␣ space ×3 0 2 15 c word starts 1 3 16–18 o d e inside 1 3 19 ␣ space 0 3 20 \0 loop test fails: stop 3 - Answer: 3 (the compiled program prints 3). The lecture's sentence, "Ada Lovelace wrote the first algorithm", gives 6.
- If the text came from
fgets, it ends in'\n'. That's not a space, so it just counts as part of the last word and the answer is unchanged. But a line that is only"\n"would count as 1 word. Strip the newline first (card 4).
Worked example 3
Predict the output: strcat, strcpy, strlen, sizeof
char a[20] = "code";
char b[] = "C";
strcat(a, "!");
printf("%d %d\n", (int) strlen(a), (int) sizeof(a));
strcpy(a, b);
printf("%s %d %d\n", a, (int) strlen(a), (int) sizeof(b));
a[1] = 'S';
printf("%s\n", a);
- Draw the boxes of
afirst.c o d e \0 0 0 …(20 boxes, zero-filled). strcat(a, "!")writes!over the 0 in box 4 and a new 0 in box 5:c o d e ! \0. Sostrlen(a)= 5.sizeof(a)is the whole array: 20. Line 1:5 20.strcpy(a, b)copiesCand its 0 into boxes 0–1:C \0 d e ! \0. Boxes 2–5 still hold the oldde!and its 0.%sprints C,strlen(a)= 1, andsizeof(b)= 2 (Cplus its 0). Line 2:C 1 2.a[1] = 'S'overwrites the 0 that ended the string, so%swalks on into the old text until the next 0 (box 5):C S d e !. Line 3:CSde!.- Output (compiled and run):
5 20 C 1 2 CSde!
The lesson: strlen measures the text, sizeof measures the array, and old characters behind a 0 are still there.
1. Slide 5, "leftover bytes". In char greeting[6] = "Hi";, boxes 3–5 are drawn as garbage ?. For an initialised array they are 0 (printed: 72 105 0 0 0 0). Garbage is real only for an uninitialised array like slide 4's char name[20];, which isn't "empty". The slide's point, that functions stop at the first 0, is right.
2. Slide 28, i <= strlen(word). On "Hi" it runs 3 passes, printing H, i and the (invisible) 0 byte, and then stops. It doesn't "keep going" past the array. It's still a bug, and < or word[i] != '\0' fixes it.
3. Slides 7–9 and 12: %s and %[^\n] aren't safe without a width. They overflow exactly like gets. Use %9s or %9[^\n] for char name[10] (size − 1), or fgets.
4. Slide 18, strcmp. It returns 0 for equal and a sign for order: negative if the first string sorts first, positive if it sorts after. So if (strcmp(a, b)) means "not equal". Always write == 0.
Classic traps: the standard ways marks are lost
| The mistake | What you see | The fix |
|---|---|---|
Array sized for the letters only: char w[8] for "Lovelace" | No room for the 0. Functions run off the end | Size ≥ length + 1: char w[9], or let C count with char w[] = "Lovelace"; |
a == b to compare text | "Not equal" even for identical text. clang: "array comparison always evaluates to false" | strcmp(a, b) == 0 |
if (strcmp(a, b)) read as "if equal" | The branches are swapped | Write == 0 explicitly |
name = "Ada"; after declaring | Compile error: arrays can't be assigned | strcpy(name, "Ada"); |
sizeof used as the length | sizeof("Ada Lovelace") array = 13 but strlen = 12. In a 20-box array, sizeof is 20 whatever the text | sizeof = boxes; strlen = characters before the 0 |
scanf("%s", name) for a full name | Only the first word lands; the rest waits in the input | scanf(" %29[^\n]", name) or fgets |
No width on %s / %[^\n] | Long input overflows: undefined behaviour (junk, crash, or "works") | Width = size − 1 |
Leftover newline after scanf("%d") | %[^\n] reads nothing (returns 0, array still junk). fgets reads just "\n" | Leading space: " %[^\n]" |
fgets keeps '\n' | strlen is one more than expected; strcmp with "Ada" fails; output breaks the line | Card 4's loop: replace the '\n' with '\0' |
strcpy/strcat into a small array | clang: "'strcpy' will always overflow; destination buffer has size 5…". The compiled program was aborted by the Mac's safety check; elsewhere it may silently corrupt | Check strlen(dest) + strlen(src) + 1 ≤ size first |
isdigit(c) == 1 | Works on a Mac (returns 1), may fail on Linux, where "true" can be another nonzero value | if (isdigit(c)) |
i <= strlen(s) | One extra pass on the 0 byte. -Wextra: "comparison of integers of different signs" | s[i] != '\0' |
Palindrome loop to len or using the array size | Every pair checked twice (harmless but slow), or letters compared against the zero tail (wrong answer) | i < len / 2 with len = strlen(word) |
gets(name) | clang: "'gets' is deprecated"; the program prints "warning: this program uses gets(), which is unsafe." | Never. fgets. |
Minimal prerequisite kit
| Fact | Where it's used |
|---|---|
A char is a small number; 'A' = 65, 'a' = 97, '0' = 48 (operators notes) | '\0' (0) is not '0' (48); strcmp's order; case matters |
Array indexing from 0, size − 1 is the last box (arrays notes) | Every string loop; the palindrome partner len − 1 − i |
Swap with a temp (Lecture 8) | Reverse in place |
| A flag: assume, then disprove (Lab 4) | Palindrome (isPal), word count (inWord) |
| int ÷ int drops the fraction | len / 2 = 2 for length 5: the middle letter is skipped |
| 0 is false, nonzero is true (operators notes) | if (strcmp(…)), if (isdigit(c)), !inWord |
What to practise
| Skill | Drill it on | Textbook backup Hanly ch. 8 | How many |
|---|---|---|---|
| What input lands where | The lesson's input widget: predict the boxes before clicking, including after "number first" | "Longer strings: concatenation and whole-line input" | 6 settings · 10 min |
| Walk to the 0: length, count, search | Rewrite strlen, the vowel count and the search by hand, then test on your own words | Self-checks for "String basics" and "Character operations" | 3 programs · 20 min |
Library calls and strlen vs sizeof | Worked example 3, then change one line and predict again | "String library functions: assignment and substrings", "String comparison" | odd-numbered · 15 min |
| Palindrome and word count, typed blind | Worked examples 1–2, then variants: ignore case; count words separated by commas | End-of-chapter programming exercises | 2 · 20 min |
| Predicting output | Predict-the-output drill (string family) | Trace the chapter's self-check snippets on paper first | 1 sheet · 10 min |
| Past-paper practice | 2024 P2 Q8 and Q11, on paper | — | 2 · 25 min |
Section titles instead of numbers, because editions renumber. In the 8th edition these should be in §8.1–8.4 and §8.6; check your copy. Skip "Arrays of pointers" and the string/number conversion sections for now: they need pointers or belong to Lecture 10. If the instructor names specific sections, go with those.
Don't start strtok, strncpy tricks, pointer-walking (while (*p)) or "top string interview questions". The deck puts pointers and strtok later, and none of it is in Lecture 9. Get the 0 rule, the three reads and the four library calls into your hands first. Type every program yourself.