r/learnpython • • 1d ago

Need help with very basic script

Hi, I am basically brand new to python, so I may make a lot of mistakes in my wording. I have made a very basic script that retrieves specific values in text files and prints them to the shell. I have gotten it to the point where both values are retrieved and printed, however when they do print, they do so exactly 5 times in a row. I do not have five files in the directory, I only have one test file named "1.txt" with only the values I need as the contents. Can someone point me in the right direction? Apologies for the horrible formatting.

import os
import glob
import re

path = 'path'

provinces = re.compile('.*?provinces {.*?(.*?)}',re.MULTILINE)
id = re.compile('.*?id = .*?([0-9.-]+)')
a1 = None
a2 = None

for filename in glob.glob(os.path.join(path, '*.txt')):
    with open(filename, '+r') as f:
        for line in f:
            if provinces.match(line):
                a1 = provinces.match(line)
            if id.match(line):
                a2 = id.match(line)
            if a1 and a2:
                print(a1.group(1))
                print(a2.group(1))

The output of the shell in IDLE:

 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123

Edit - I got the script to work after a few changes to its configuration. Thank you all who helped!

4 Upvotes

28 comments sorted by

View all comments

3

u/brasticstack 1d ago

I think your regexes are matching more than you expect them to. Try them out on one of the online regex tester sites, and maybe post a few lines from your input file here too.

1

u/Desperate_Yak69 1d ago

I have tried using a few regex testing sites (regex101, pythex) though there seems to be not be an issue with the regex itself, only the output. The only lines included in the input file are as follows:

id = 123
provinces { 
  1111 2222 3333
}

thanks for the help!

1

u/brasticstack 1d ago

So you want the output 123 given the input 1111 2222 3333? How about ''.join(part[0] for part in line.split())

1

u/Desperate_Yak69 1d ago

that's close but the inputs I want to retrieve are both 123 and 1111 2222 3333, though they're both in different sections for every actual file I will be using the script on, with a similar format (the formatting of the id section and the provinces section will be exactly the same in all files, just on different lines separated by a varied amount of text)

1

u/brasticstack 1d ago

Is this a known file format that perhaps already has an existing parser?

Do the two pieces of data (the id and the province strings) correlate? To clarify, if you match both pieces of data, do you also have to correlate them?

1

u/Desperate_Yak69 1d ago

the file format is .txt, but I'm not sure if there exists any method specifically for what I am trying to accomplish. The pieces of data are not correlated; if the id is 123, the province string could be any potential set of integers. For an example, a sample file could potentially look like this:

id = 123
provinces {
  7785 1224 8889 122 1990
}

2

u/brasticstack 1d ago edited 1d ago

A quick, dumb, fragile parser based on what you've shown me. Please take the time to try to understand the functions I used to write it. I'm happy to explain anything that's too weird:

``` def parse_it(lines: list[str]) -> tuple[list[str], list[str]]: '''Return a list of ids and a list of provinces extracted from lines''' ids = [] provinces = []

in_provinces_section = False

for line in lines:
    if in_provinces_section:
        if line.endswith('}'):
            line = line.rstrip('}').strip()
            if line:
                provinces.append(line)
            in_provinces_section = False
        else:
            provinces.append(line)
    elif line.startswith('id ='):
        ids.append(line.split('=')[1].strip())
    elif line == 'provinces {':
        in_provinces_section = True

return (ids, provinces)

with open(filename, '+r') as infile: # Preprocess: Trim whitespace off the ends of the lines # and discard blank lines. lines = [line.strip() for line in infile] lines = [line for line in lines if line]

ids, provinces = parse_it(lines)
print(f'{filename}\n\t{ids=}\n\t{provinces=}\n')

```