This project has been created as part of the 42 curriculum by jiylee.
- Helping me understand an important concept in C programming: static variables.
- This activity is about programming a function that returns a line read from a file descriptor. Its version is 14.2 including the Bonus part.
-
- A function that returns a line read from a file descriptor, reading one line at a time with each successive call.
- An enhanced version that can handle multiple file descriptors simultaneously, allowing interleaved reading from different sources without losing track of each descriptor's reading position. This is implemented using only a single static variable.
git clone <repository-url> working_directory
cd working_directorySince this project has no Makefile, compile directly with cc and the -D flag to define BUFFER_SIZE.
Mandatory:
cc -Wall -Wextra -Werror -D BUFFER_SIZE=42 main.c get_next_line.c get_next_line_utils.c -o programBonus:
cc -Wall -Wextra -Werror -D BUFFER_SIZE=42 main.c get_next_line_bonus.c get_next_line_utils_bonus.c -o programYou can change the value of BUFFER_SIZE to any positive integer to test different read sizes.
Type ./program to execute.
#include "get_next_line.h"
#include <fcntl.h>
int main(void)
{
int fd;
char *line;
fd = open("example.txt", O_RDONLY);
while ((line = get_next_line(fd)) != NULL)
{
printf("%s", line);
free(line);
}
close(fd);
return (0);
}Compile with:
cc -Wall -Wextra -Werror -D BUFFER_SIZE=42 main.c get_next_line.c get_next_line_utils.c -o programExecute the program with:
./program-
-
- I used only Opus 4.6 of Claude.
-
- Getting hints of how logics should flow when I got stuck.
- Getting detailed knowledge of everything like knowing full-name of each title.
- Getting evaluation of my original source-code and refactored one, to choose which one should survive.
- Getting to know why this code must be like this, like each intention of some codes, which cases I should use which code, as the right code in the right place.
- Writing this README.md, then I edited it.
-
The read system call does not care about newlines. When we call read(fd, buf, BUFFER_SIZE), it simply reads up to BUFFER_SIZE bytes regardless of where \n appears. This means a single read call may contain more data than one line, or a line may span across multiple read calls. We need somewhere to store the leftover data between function calls. A local variable would be destroyed when the function returns, so a static variable is the natural choice, as it retains its value across multiple invocations of the same function.
get_next_lineis called with a file descriptorfd- It reads from
fdinto a temporary buffer ofBUFFER_SIZEbytes using thereadsystem call - The read content is appended to the static variable (the "backup") that persists across calls
- Reading repeats until a newline
\nis found in the backup orreadreturns 0 (EOF)
Why loop until \n is found?
A single read call with a small BUFFER_SIZE may not capture a full line. For example, if BUFFER_SIZE is 4 and the first line is "Hello\n", it takes two reads ("Hell" then "o\n") to reach the newline. Conversely, if BUFFER_SIZE is large, one read may capture multiple lines at once, which is why the leftover must be saved.
- Once the backup contains a newline, it is split at the first
\n - Everything up to and including
\nis returned as the current line - Everything after
\nremains in the backup for the next call - If there is no newline and EOF is reached, the remaining backup is returned as the last line
Why include \n in the returned line?
The subject requires it. This also allows the caller to distinguish between a mid-file line (ends with \n) and the final line of a file that has no trailing newline (no \n at the end).
- Instead of a single
static char *backup, astatic char *backup[1024]array is used - Each index corresponds to a file descriptor, so
backup[fd]stores the leftover data for that specificfd - This allows interleaved calls like
get_next_line(3),get_next_line(4),get_next_line(3)without losing any reading position
Why an array indexed by fd?
File descriptors are non-negative integers assigned sequentially by the OS, making them perfect natural indices for an array. Compared to alternatives like a linked list of (fd, backup) pairs, an array provides O(1) direct access with no search overhead. {O(1) access to an array is possible because of the combination of three factors: random access characteristics of RAM + contiguous memory layout + address calculation (addition).} The trade-off is that it allocates space for 1024 pointers regardless of how many are used, but since each element is just a pointer (8 bytes), the total memory cost is negligible. In reality, even if only three file types are used, 1,024 pointers are allocated. However, 8KB is a negligible size by modern computer standards. That's why the the text used the term "negligible". On the other hand, a linked list saves memory because it only creates nodes for the three file types actually used, but each search takes O(n) time.
Array (backup[1024]) |
Linked List | |
|---|---|---|
| Lookup speed | O(1) — direct access | O(n) — traversal required |
| Memory usage | ~8KB fixed (mostly empty) | Only as much as the number of active fds |
| Implementation complexity | Simple | Requires node creation, deletion, and search logic |
| Conclusion | Fast ✅ slight memory waste | Memory-efficient ✅ slower |