Python用re正则化模块在字符串查找特定字符串

LazyCat

2017-09-29

实验需要，在一个含有几亿个字符的txt文件中查找特定的字符串，首先用re模块进行查找

from time import clock
import re
start=clock()
label_file = open("/home/ying/data/google_streetview_train_test1/label.txt")
label_str = label_file.read()
label_file.close()
filename = "2_0_pitch_95_yaw_95_lat_41.8975137_lng_-87.6268723.jpg"
start=clock()
for match in re.finditer(filename, label_str):
s = match.start()
e = match.end()
print(s)
print(e)
end=clock()
print(end-start)

re.finditer(filename, label_str)可以在label_str中查找filename的位置，s=match.start()返回字符串开始的索引，e=match.end()，返回字符串结束的索引。程序运行的结果是

304091635
304091689
304096479
304096533
1.003844

耗时1s左右

同样的，由于txt文件中为一行一行的数据，可以用readlines进行遍历读取比较，程序如下

from time import clock
start=clock()
data_label="/home/ying/data/google_streetview_train_test1/label.txt"
filename = "2_0_pitch_95_yaw_95_lat_41.8975137_lng_-87.6268723.jpg"
file = open(data_label)
lines = file.readlines()
print(len(lines))
for line in lines:
cls = line.split()
fn = cls.pop(0)
if fn==filename:
break
end=clock()
print(end-start)

运行结果如下：

1
3.335657

可见耗时有3s多，用正则化模块要快的多

另外，由于label_str中存在1.2_0_pitch_95_yaw_95_lat_41.8975137_lng_-87.6268723.jpg，所以用re模块寻找时会返回两个结果，而用逐行读取的方式则返回一个值

python字符串正则化 python label

安科网

Python用re正则化模块在字符串查找特定字符串

LazyCat

LazyCat

相关推荐

Python判断字符串以什么开始

python第二天

python

Python格式化字符串(f,F,format,%)

Python用户交互、格式化输出及运算符

Python字符串前缀u、r、b、f含义

python字符串和编码

python之字符串split和rsplit的方法

字符串的查找删除

4_2 json字符串转Python对象

Python 正则表达式

Golang中的Unicode与字符串示例详解

Python 字符串函数_删除

1.Python实现字符串反转的几种方法

python基本数据类型；关于字符串格式化的不出

python字符串的表示方式

python实现将固定格式的字符串调整为字典的格式，用于爬虫爬取数据时快速添加请求数据

python端口IP字符串是否合法

python生成随机数、随机字符串

字符串&列表&元组&字典之间互转

LazyCat