CS U111 · Lecture 9 · Lesson: first time through
Strings, from the beginning
One idea at a time, with the reasoning before the syntax. Allow about 45 minutes, and try each "check yourself" before you open it. The short version is on the notes page.
If Lecture 8 made sense, most of this lecture is already yours. A string in C is an array of char, so indexing, loops and "C never checks the bounds" all carry over unchanged. The genuinely new material is small: one rule about how a string ends, a few ways of reading a line from the keyboard (where most of the bugs live), and a short list of library functions. Strings are new to everyone in the batch. School CS in Python or C++ hides all of this behind a ready-made string type, and C doesn't have one.
Scope: Lecture 9 (Strings in C), whose reading is Hanly & Koffman ch. 8. The deck says Lecture 10 goes deeper into the library, and that passing strings to functions and strtok wait until after pointers. Neither appears here.
1 · Text is just characters, and characters are numbers
You met this in the operators lesson: a char is a one-byte whole number, and the quotes in 'A' are a way of writing the number 65 so you don't have to remember it. Printing with %c shows the character; printing with %d shows the number.
So a piece of text such as Hi is two small numbers in a row: 72 and 105. You already know how to keep values in a row: an array. That's the whole trick. C has no special "string" type; it uses a char array.
That leaves one problem. An int array of marks has a size you track yourself. Text changes length all the time: a name array might hold "Ada" today and "Lovelace" tomorrow. How does a function like printf know where the text stops? That's the one new rule.
2 · The one rule: a string ends at '\0'
After the last real character, C stores a byte whose value is 0. It's written '\0' (backslash-zero) and called the null terminator. It's not the digit '0', which is 48. Every string function reads characters one after another until it meets that 0. The array's size doesn't matter to them at all; the 0 does.
When you write text in double quotes, C adds the 0 for you. "Hi" is three bytes: 72, 105, 0.
char greeting[6] = "Hi";: the bytes as stored (checked by printing each one with %d: 72 105 0 0 0 0).Two things to notice in the picture.
- The size must leave room for the 0. "Hi" needs 3 bytes, not 2. Declaring
char greeting[2] = "Hi";leaves no room for the 0, so it isn't a string any more. clang warns:initializer-string for character array is too long, array size is 2 but initializer has size 3 (including the null terminating character). - The unused boxes are zeros, not garbage. The lecture slide draws boxes 3–5 as
?. For an array that's initialised like this one, C fills every box you didn't give a value with 0. It's the same rule you saw in Lecture 8, whereint part[5] = {1, 2};became1 2 0 0 0. Either way the slide's real point stands: string functions stop at box 2 and never look further.
Garbage is real for an array you declare without initialising:
char name[6]; inside a function: whatever bytes were already there. There may not even be a 0 anywhere.The slide that calls char name[20]; "empty" means "nothing useful in it yet". Don't print it, or pass it to strlen, until something has put a string into it; with no 0 inside, those functions just keep reading.
Check yourself: what's the smallest array that can hold the word Lovelace as a string? What does char w[] = "Lovelace"; get as its size?
8 letters plus the 0, so 9. With empty brackets C counts for you and gets the same answer: sizeof(w) is 9.
3 · Printing: %s walks until it meets the 0
printf("%s", name) means: start at box 0 and print characters until you reach a 0 byte. %c prints exactly one character, such as name[0]. Because a string is an array, you can change one box: after greeting[0] = 'B'; the same array prints as Bi.
The 0 is a marker, and markers can be moved. That explains a result that looks like magic at first:
char a[20] = "code";
strcpy(a, "C"); /* a is now C, 0, d, e, 0, 0 ... */
printf("%s\n", a); /* C */
a[1] = 'S'; /* overwrite the 0 that ended the string */
printf("%s\n", a); /* CSde (the old letters are back) */
strcpy copied two bytes, C and a 0, into the front of the array. The old de is still sitting behind the new 0. Replace that 0 with a letter and %s keeps walking into the old text. (Compiled and run: it prints C, then CSde.)
Check yourself: char s[] = "Ada"; then s[1] = '\0';. What does printf("%s", s) print, and what is strlen(s)?
It prints A. The walk starts at box 0, prints A, meets a 0 in box 1 and stops. strlen(s) is 1. The a in box 2 is still in memory; nothing reads it.
4 · Reading from the keyboard
This is where most string bugs happen, and it's easier with one picture in your head. When you type Ada Lovelace and press Enter, those characters, including an invisible newline character '\n' for the Enter key, go into a queue. Each scanf or fgets takes characters from the front of that queue. Whatever it doesn't take stays there for the next read.
The four ways the lecture shows
scanf("%s", name)skips any spaces or newlines first, then takes characters up to the next space or newline. Ada Lovelace givesAda; " Lovelace" and the'\n'stay in the queue. There's no&beforename: an array's name already says where it lives in memory (why, exactly, comes with pointers).scanf("%[^\n]", name)takes every character that is not a newline, so the whole line, spaces included. The'\n'itself stays in the queue. The brackets hold a scanset:%[^,]stops at a comma,%[0-9]takes only digits.scanf(" %[^\n]", name), with a space before the%. In ascanfformat, a space means "skip any spaces and newlines here". You need it when an earlierscanf("%d", …)left its'\n'in the queue.fgets(name, 10, stdin)takes at most 9 characters (it keeps one box for the 0), and stops early after a newline, which it keeps. The 10 is the array's size, so it can't overflow.
And one never to use: gets(name). It reads a whole line but is never told the array's size, so a long line writes past the end. It was removed from the C standard in 2011. On a Mac, clang warns that 'gets' is deprecated, and the running program prints warning: this program uses gets(), which is unsafe.
Try it: what actually lands in name
The program below has char name[10];. Type a line, pick a way of reading it, and see the bytes that land in the array and what's left in the queue. Try a long line to see an overflow, then pick the version with a width. Tick the box to put scanf("%d", &age); first, as in the lecture's "practical gotcha" slide.
The widget follows the same rules as the real C library. Its results were checked against compiled programs for 144 combinations of line, method and "number first" (12 lines × 6 methods × with or without the number). Boxes marked ? were never written, so they hold whatever was there before. Red dashed boxes are past the end of name[10]: that's undefined behaviour, and a real program may crash, print junk or appear to work.
Four things the widget makes visible:
%sstops at the first space. With Ada Lovelace, onlyAdalands; "␣Lovelace⏎" waits in the queue for the next read.- No width means no limit. Type Lovelace1234 with
%s: 12 letters plus the 0 need 13 boxes, and the array has 10.%9stakes 9 letters, adds the 0 in box 9, and leaves the rest in the queue. The width is always the array size minus one. The lecture presents%[^\n]as the safe way to read a line, but without a width it can overflow exactly likegets.%9[^\n]can't. - The leftover newline. Tick "number first" and choose
%[^\n].%dreads 25 and leaves the'\n'.%[^\n]doesn't skip anything, so the very first character it sees is the newline it must stop at. It reads nothing,scanfreturns 0, andnameis untouched. Your program then prints whatever junk was in it. The lecture shows the fix (the leading space) but not this failure, and it's the one you'll meet. fgetskeeps the newline. Short lines arrive asA d a \n \0. After a%d,fgetsreads only the leftover newline and returns straight away, sonameis just"\n". (The notes page shows the one-loop fix to remove it.)
Check yourself: char city[8];. Which format reads one word safely?
scanf("%7s", city). That's 7 characters plus the 0 = 8 boxes. The width counts characters, not bytes, so it's always size − 1.
Check yourself: the input is 30 Enter New Delhi Enter. The program runs scanf("%d", &n); fgets(city, 20, stdin);. What is in city?
Just "\n": a newline and a 0. fgets stops after the first newline, and the first thing in the queue is the one %d left behind. "New Delhi" is still waiting. Fixes: use scanf(" %19[^\n]", city) instead, or read and throw away the rest of the number's line before calling fgets.
5 · Walking a string: stop at the 0
Almost every string loop has the same shape as a Lecture 8 traversal, with one change: the stopping test is the 0, not a size.
for (int i = 0; word[i] != '\0'; i++) {
/* do something with word[i] */
}
Counting the steps of that loop is exactly what strlen does: "Lovelace" gives 8, and the 0 isn't counted. The lecture builds it by hand:
int len = 0;
while (word[len] != '\0')
len++;
The off-by-one, precisely
The lecture's warning slide uses for (int i = 0; i <= strlen(word); i++) on "Hi". It's a real bug, but the slide overstates it. strlen is 2, so i runs 0, 1, 2. That's 3 passes, and pass 3 prints the 0 byte, which is invisible. Then i becomes 3, the test fails, and the loop stops. It doesn't "keep going" past the array. (Compiled: it printed bytes 72, 105, 0 and made 3 passes.) The fix is the lecture's: <, or better, the word[i] != '\0' test above. -Wall alone says nothing here. -Wextra adds comparison of integers of different signs: 'int' and 'unsigned long', because strlen returns an unsigned type. That's also why the lecture writes printf("%d", (int) strlen(name)).
Check yourself: how many vowels does the lecture's loop count in "Programming"?
3: o, a, i. Walk it: P r o g r a m m i n g. (Compiled answer: 3.)
6 · The string.h toolkit
Add #include <string.h>. Each of these functions is one of the loops above that someone has already written for you:
strlen(s): the number of characters before the 0.strcpy(dest, src): copiessrc, including its 0, intodest. It's how you "assign" a string:dest = src;doesn't compile for arrays.strcat(dest, src): finds the 0 at the end ofdestand copiessrcfrom there.strcmp(a, b): compares letter by letter. It returns 0 when the texts are equal, a negative number whenacomes first in character-code order, and a positive number whenacomes after.
Why == doesn't compare text
a == b on two arrays compares where they live, not what they hold. Two separate arrays never live in the same place, so it's always false, even when both hold "Ada". clang says so: warning: array comparison always evaluates to false [-Wtautological-compare]. Use strcmp.
Now the trap the slide doesn't mention. Because "equal" is 0, and 0 means false in an if:
if (strcmp(a, b)) /* true when they are DIFFERENT */
if (strcmp(a, b) == 0) /* true when they are the same */
Always write the == 0. It reads the way you mean it.
Check yourself: is strcmp("apple", "Apple") positive, negative or zero?
Positive. The first letters differ: 'a' is 97 and 'A' is 65. 97 comes after 65, so "apple" sorts after "Apple". Lowercase comes after uppercase in character codes.
Nothing checks the destination size
strcpy and strcat write as many bytes as the source needs, however small dest is: the same danger as gets. For char small[5]; strcpy(small, "Lovelace");, clang even warns 'strcpy' will always overflow; destination buffer has size 5, but the source string has length 9 (including NUL byte). The compiled program was stopped by the Mac's built-in safety check. Elsewhere it might print junk or seem to work. Before a copy or a join, check the total fits: strlen(dest) + strlen(src) + 1 must be at most the array's size.
7 · One character at a time: ctype.h
#include <ctype.h> gives questions and conversions for a single character: isalpha(c), isdigit(c), toupper(c), tolower(c). They don't loop; you write the loop. toupper leaves anything that isn't a lowercase letter alone, so toupper('3') is still '3'.
One portability trap: the is… functions return nonzero for true, not necessarily 1. A Mac happens to return 1, but the C standard only promises "nonzero", and Linux lab machines may return a different nonzero value. So write if (isdigit(c)), never if (isdigit(c) == 1).
8 · Changing a string in place
Because a string is an array, you can rewrite its boxes where they are. The lecture's uppercase loop is just word[i] = toupper(word[i]); inside the standard walk.
Reversing is Lecture 8's swap-the-ends idea: swap box i with box len − 1 − i, only up to the middle (i < len / 2). Go all the way and every pair is swapped twice, which puts the string back. The lecture's example, "Ada", reverses to "adA". The capital moves to the end, which is a nice check that the swap really happened.
The palindrome activity uses the same pairing without swapping: compare box i with box len − 1 − i and stop at the first mismatch. The full solution with a trace is worked example 1 on the notes page.
Check yourself: for a word of length 5, which pairs does the palindrome loop compare? What about the middle letter?
len / 2 is 2 (integer division), so i is 0 and 1: pairs (0, 4) and (1, 3). The middle box 2 has no partner. It always matches itself, so it's never checked. For length 4, the pairs are (0, 3) and (1, 2), with no middle.
9 · Finding words: remember where you are
To count words, walk the sentence with one extra variable, a flag inWord that remembers whether the previous character was part of a word. A word starts at the moment you go from a space into a non-space: count it then, and set the flag. A space clears the flag. Because you count starts rather than spaces, double spaces and leading spaces don't create extra words. The full trace on a sentence with double spaces is worked example 2 on the notes page.
Searching by hand is Lecture 8's linear search on characters: the lecture finds 'L' in "Ada Lovelace" at index 4 (A0 d1 a2 space3 L4). The library versions, strchr and strstr, return a pointer, so they wait until after pointers.
10 · You're ready: what to do next
- Revise from the notes page: rule cards, three worked examples at exam level, and the trap table.
- Type the palindrome and word-count programs yourself and run them on your own inputs, including one with double spaces.
- Predict-the-output practice: the functions & arrays drill.
"I'm learning C strings (Hanly & Koffman ch. 8, no pointers yet). Give me five short programs that read input with scanf %s, %[^\n] or fgets, with the exact input typed. Let me predict what ends up in the array, then show me the byte-by-byte answer, including '\n' and '\0'."
"Quiz me on strcmp: give me 8 pairs of strings, I'll say whether strcmp returns negative, zero or positive. Then show me an if statement using strcmp that has a bug, and let me find it."
"Here is my palindrome checker [paste it]. Don't rewrite it. Ask me questions that would make me find any bug myself, especially around the loop bound and the middle letter."
Type every program yourself. Lab evaluations test your hands, not the AI's.