Showing posts with label python. Show all posts
Showing posts with label python. Show all posts

Wednesday, May 27, 2009

Python module of the week

Someone has been doing a tutorial of a python module per week for quite some time

Wednesday, March 25, 2009

Read multiple newline types in csv

One really convenient module I use all the time is the csv module. You would think that parsing csv by yourself is easy, i.e., if you've never tried it, but it's really rather tricky.

People who use excel (sneer) often need to send me csv-formatted spreadsheets, and quite often the formatting is quite messed up. Often the line-endings are bizarre and frightening, and unrecognizable.

The solution is to add a 'U' flag to the open() call to enable universal newline support.

For instance, if you're opening a file for reading:

open('somefile.csv', 'rU')

will not care what the newline format is. Somehow it manages it all without me having to think about it. My thoughts are expensive.

Tuesday, March 24, 2009

Monday, February 16, 2009

Disposable variables

Ever have a bunch of variables you get returned in a tuple, and you only want some of them? So you assign them to some useless variable like 'nothing' or 'deleteme' and hope that the optimizer will instantly garbage collect them away?

(user, user_type, ingorethis) = somefunction()

The correct way to do this is to use the _ variable. (underscore). Underscore is equivalent to perl's $_ variable: it contains the results of the last evaluated expression. This is sometimes useful if you're using the python shell as a calculator. So, it will instantly be overwritten.

(user, user_type, _) = somefunction()

Saturday, December 20, 2008

make python understand ~ for home directories

Python does not automatically expand ~ for user home directories. There is a function in os.path that does that for you:

mypath = os.path.expanduser('~/somepath')
mypath == '/home/nick/somepath'

any and all

Python has two builtin functions for evaluating iterables for truth/falseness: any() and all(). any() returns true if any member of the iterable evaluates to True, while all returns True only if every single member of the iterable evaluates to True.

>>> any([True, False, False])
True
>>> any([False, 0])
False
>>> all([True, 0])
False
>>> all([1])
True
>>> all([1, 1, True])
True

There are identically named subquery operators in sql, however they are only really useful for correlated subqueries which are poorly optimized and generally to be avoided in the current versions of mysql.

Friday, December 12, 2008

Parse any date easily

Use dateutil.parser.parse, possibly with the fuzzy flag set to True.

>>> from dateutil.parser import parse
>>> parse("Jan. 1st, 1991", fuzzy=True)
datetime.datetime(1991, 1, 1, 0, 0)

This is a huge boon over the tedious and underfeatured time.strptime and similar. Unhelpfully, the documentation is non-existent.

Thursday, December 11, 2008

See how python compiles some code

use the dis module to get the bytecode of any python callable.

Example:

import dis
dis.dis(lambda: x is True)



  1           0 LOAD_GLOBAL              0 (var)
              3 LOAD_GLOBAL              1 (True)
              6 COMPARE_OP               8 (is)
              9 RETURN_VALUE        

This is another way to get the efficiency of different parts of code

Sunday, December 7, 2008

See how python encodes a string internally

Fun fact: you can see how python encodes a string internally by asking it to encode a unicode object to the 'unicode_internal' codec.

Example:


>>> u'\N{SNOWMAN}'.encode('utf16')
'\xfe\xff&\x03'
>>> u'\N{SNOWMAN}'.encode('unicode_internal')
'&\x03'
As you can see, it's UTF16 without a BOM.

Fun stuff with unicode in python

Fun fact: you can interpolate a unicode character into a unicode string literal by using \N{name of character}

Example:

>>> print u'this sign is \N{VULGAR FRACTION ONE HALF} off'.encode('utf8')
this sign is ½ off

You can tell the name of any given character by using the unicodedata module

>>> import unicodedata
>>> unicodedata.name(u'\xbd')
'VULGAR FRACTION ONE HALF'

The name of the famous snowman character is simply enough 'SNOWMAN.

>>> unicodedata.name(unichr(9731))
'SNOWMAN'


So you can represent this character in python simply by u'\N{SNOWMAN}'

Tuesday, December 2, 2008

Keep a series of python objects in a shelf

You can store a dict of pickled objects in a "shelf", a dbm-based hash table stored to disk, using the shelve module.

Sunday, November 30, 2008

Web-based auto-generated self-documentation

Fire up pydoc -p 8080, then browse to http://localhost:8080

It will generate a browseable directory of documentation for every module you have installed that's visible on sys.path .

Though the layout is hideous.

I'm not aware of another way to see every module you have installed.

Friday, November 28, 2008

Never use a [] or {} as a default argument.

I've done this a couple times.

When you make [] a default argument, it will be the same list on every subsequent invocation, with the same id(). If you append to it, what you've appended will be still be there for later invocations of the function.

Example:

def functionx(mylist = []):
    mylist.append("F")
    print mylist

for i in range(4): functionx()

The output of this program should be:
F
F, F
F, F, F
F, F, F, F

Needless to say, this is really, really stupid. The reason is because the default arguments get defined at definition time, and any mutable type will persist. This is also why somedict.setdefault({},...) will also mess up, any why defaultdict instead uses a callable.

So, you should only use constant values as default parameters, and never a mutable type. I've made this mistake a couple times.

I've seen some obtuse lunkheads actually saying it's better this way. It seems to be a rule that all language designers eventually start rationalizing their mistakes and the corners they cut (another example: see this interview with HÃ¥kon Lie , as he is unable to admit there could be any flaw in the design of CSS) and pointlessly defend bad decisions in the work they've invested so much time that no one who approached the language with fresh eyes would ever, ever defend.

I've also seen some people saying that you can use this to simulate static variables. To which I say: so, give us proper static variables then, and don't plant your language full of landmines sure to screw people up when they're not expecting it.

Wednesday, November 19, 2008

More complete regex support coming in 2.7?

Looking through the python issue tracker, looks like they're working on almost my exact wishlist of missing features in python's re module.

Including:
  • possessive quantifiers, atomic grouping
  • (?flag:...), (?-flag:...) support, (?flag) scoped correctly
  • variable length lookbehinds
So all of this will be ready for 2.7, right? right?

Monday, November 17, 2008

Get subpattern matches balanced even if they're optional

This will give you a mysterious "unbalanced grouping" error:

re.sub('.(.)?)', r'aaa\1')

Because the parenthetical grouping is optional, trying to access it via the \1 backreference will result in an error.

The solution is to replace the ? operator with a null alternative in the paren:

re.sub('.(|.)', r'aaa\1')

Which gives the grouping the option of containing nothing, thereby maintaining the existence of a backreference (though the backreference will of course contain an empty string, which is what you would actually expect).

Wednesday, November 12, 2008

Nested list comprehensions

Fun fact: you can nest list comprehensions.

This syntax looks like:

[(i, j) for i in range(4) for j in range(i)]

Which is exactly equivalent to:

x = []
for i in range(4):
for j in range(i):
x.append((i, j))

Monday, November 10, 2008

The hidden debug flag in the re module

The compile function of the re module, as I'm sure you already know, takes a second parameter of flags, such as re.I or re.X.

There is an undocumented flag, however, that will print out the parse tree of the regex. This is the re.DEBUG flag.

Here's some trivial output:

>>> re.compile(u'ell2\u1f4c', re.DEBUG)
literal 101
literal 108
literal 108
literal 50
literal 8012
The 'literal' statements are followed by the ascii dec values of the characters in the pattern. I threw in a unicode character and you can see the unicode code point as the last item.

This will also tell you whether the pattern is being compiled for the first time or if it is being pulled out of the re module's cache.

Thursday, November 6, 2008

Make a more efficient list of integers

If you have a list that's all of the same type, and that type is numeric, use the array module for more efficiency.

Keep a list always sorted

Use the bisect module

Tuesday, November 4, 2008