Saturday, November 20, 2004

Performance counter patch for Linux

Some of you have already installed this, but if not, here's the link to the perfctr patch. If you have any other links to performance-counter related stuff, please add it in comments to this post.


2 Comments:

Blogger Emery said...

Here are some more performance counter related tools. Not sure which ones are good... Brink and Abyss, PerfSuite, and OProfile.

11/24/2004 10:29 AM  
Blogger Yi said...

After spending much time playing perfex/lperfex, I have finally discovered the following (for Pentium 4 only!):

1) lperfex does NOT support P4 at all

2) it's very difficult to understand how to use perfex - its documentation is really poor

3) after looking at the Intel P4 manual and decoding the available perfex usage examples (none of which is about L2 cache miss), I have worked out two potential ways to get L2 cache miss counts:

#1:
There is an example in perfex.c to count L1 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000001 --p4pmv=0x1 some_program

Change it to the following to count L2 cache read misses:
perfex -e 0x0003B000/0x12000204@0x8000000C --p4pe=0x01000002 --p4pmv=0x1 some_program

On the P4 manual, this is a replay event (god knows what it is) and the replay metric is called 2ndL_cache_load_miss_retired, which has a footnote: 2nd-level misses retired does not count all 2nd-level misses. It only includes those references that are found to be misses by the fast detection logic and not those that are later found to be misses. [IA-32 Manual Vol. 3 A-35]

#2:
I finally got to know how to set the magic numbers (these numbers are not anywhere else on the internet):

perfex -e 0x0003f000/0x18020004@0x80000000 some_program
for RD_2ndL_MISS: read 2nd level cache miss (includes load and RFO).

perfex -e 0x0003f000/0x18000e04@0x80000000 some_program
for RD_2ndL_HITS + RD_2ndL_HITE + RD_2ndL_HITM: read 2nd level cache hit Shared/Exclusive/Modified (includes load and RFO).

I tested them, and it seems they are truely for the L2 cache. However, the number reports by #1 is always 3 ~ 5 times larger than that reported by #2. They are consistent with themselves as I change the test program's working set.

12/05/2004 4:23 PM  

Post a Comment

<< Home