Wednesday, March 23, 2016

Plugging Gaps


So the previous set of compiler directives singlehandedly fixed a wide-ranging set of errors when I tried to build the Linux code. But there were plenty more in there -- here are some others I found interesting to track down.

The porting code I had brought over included a couple already, such as likely and unlikely, which are hints to the compiler about the path it would do better to optimize for:
#define likely(x)       __builtin_expect(!!(x), 1)
#define unlikely(x)     __builtin_expect(!!(x), 0)

There were a couple more obvious ones, such as __init and __exit used to designate the startup and shutdown functions for a Linux module. Since this code is all going to be one "module" (e.g. kext) on OS X and it comes with its own API for startup and shutdown, I just disabled those:
#define __init
#define __exit

Another simple copy from the Linux source was BUILD_BUG_ON, which aims to cause a compiler error:
#define BUILD_BUG_ON(condition) \
       ((void)sizeof(char[1 - 2*!!(condition ? 1 : 0)]))
While a little obscure, this relies on the fact that the calculated value will either be compiled out of existence or cause an error:
(void)sizeof(char[0])   // Compiled away
(void)sizeof(char[-1])  // Compiler error

The trickier ones were more embedded into the way things work in Linux, such as debugging. WARN_ON prints a message when the provided expression evaluates to true, and returns the value for use in an if statement. So the usage looks like this:
if(WARN_ON(foo == 3)) {
    return -EINVAL;
}
Well, WARN_ON causes an Oops, similar to a panic on Linux except it doesn't stop the machine dead. OS X doesn't have an apparent analog. Further, the Linux implementation quickly descends into large functions in printk.c, which also doesn't have an OS X equivalent. Instead, I went with the workaround of defining a global function to save a bug to the log. And because it's not expected to happen very often, I went with a fairly straightforward one and then saved the message to the log with IOLog. The only tricky bit is that is uses a printf-style string, so I had to handle variable arguments:
static void porting_print_warning(char *fmt, ...) {
    char buffer[200] = "AppleIntelWiFiMVM FATAL ";
    char *remainder = &buffer[24];
    va_list args;
    va_start(args, fmt);
    vsnprintf(remainder, 176, fmt, args);
    va_end(args);
    IOLog(buffer);
}

That made the WARN_ON implementation an easy copy from the Linux source, just replacing the function to call to log the actual warning:
#define WARN_ON(condition) ({                                 \
    int __ret_warn_on = !!(condition);                        \
    if (unlikely(__ret_warn_on))                              \
        porting_print_warning("at %s:%d/%s()!\n",             \
                              __FILE__, __LINE__, __func__);  \
    unlikely(__ret_warn_on);                                  \
})

The similar BUG and BUG_ON are a little trickier, because they're supposed to print a message and then panic the kernel. Well, the panic part is OK, but I think the message part is the problem -- I believe IOLog saves to an in-memory buffer, and something else flushes the buffer to the system log file. And I expect when the line after IOLog is panic(); probably the buffer never gets flushed to the log file. I've read that OS X saves the panic data to nvram in order to display it on the next boot, exactly because nothing can be reliably saved to disk in the event of a panic. So here's my implementation, which probably doesn't really work as desired (a note for future investigation):
#define BUG() do {                             \
    IOLog("BUG: failure at %s:%d/%s()!\n",     \
                __FILE__, __LINE__, __func__); \
    panic("BUG!");                             \
} while (0)

#define BUG_ON(condition) do {      \
    if (unlikely(condition)) BUG(); \
} while (0)

Now, you may be curious about the specific syntax used in the last couple examples (I know I was):
({...; foo;})

do {...} while (0)

The first, it seems, is the way you write some code that does something and then return foo from a macro that will be substituted into the condition of an if statement (shown here after substitution):
if(({...; foo;})) {...}
But also works (albeit suboptimally) as the result of the if statement instead of the condition:
if(foo) ({...; foo;});

The other is how you write a macro that returns nothing and may be used as the result of an if block with or without curly braces:
if(...)
    do {...} while (0);
or
if(...) {
    do {...} while (0);
}
Without the do/while, you might end up with something like this after substitution:
if(...)
    IOLog("BUG: failure at %s:%d/%s()!\n", ...);
panic("BUG!");

Ah, code that sometimes logs but always panics. That would deserve its own Oops.

Compiler Directives

<< Prev: Xcode vs. AppCode       Next: Plugging Gaps >>

So there I was, faced with some inexplicable snippet of Linux code. I googled it. I landed at http://lxr.free-electrons.com/, which has got to be one of the single most useful resources for porting Linux code, hands down.

Whatever I was looking at was a compiler directive. Probably it was something like __maybe_unused. I followed the link on that page to compiler-gcc.h. And what do you know? Right above __maybe_unused is __packed. Remember __packed? I remember. It took me at least an hour to eliminate every occurrence of __packed.

But wait. That was way back when I thought __packed was complete magic. Now it turns out it's just a macro for an annotation that the compiler picks up.

But that's for GCC, and OS X has switched away from GCC. I'm pretty sure that's true -- I mean, I had to use lldb instead of gdb, and the project screen has a bunch of settings for Apple LLVM 7.0. But, I thought I heard that gcc and some other compiler forked and merged... maybe that was LLVM? (No.) Maybe they're pretty close after all? (Yes. Despite being different projects, GCC compatibility appears to be a goal of clang, the C/C++ compiler "front end" to LLVM.)

Bottom line, could I just snarf the definition of __packed instead of rewriting all that code? Yes I could. Yes I should, since that would make it a lot easier to update to the latest driver code later.

And while __packed (and __aligned) were the most obvious, there were a lot of others:
#define __packed             __attribute__((packed))
#define __aligned(x)         __attribute__((aligned(x)))
#define __printf(a, b)       __attribute__((format(printf, a, b)))
#define __attribute_const__  __attribute__((__const__))
#define __maybe_unused       __attribute__((unused))
#define __bitwise__          __attribute__((bitwise))
#define __must_check         __attribute__((warn_unused_result))

Most of those I had previously #defined to nothing, so it was nice that they would work more as intended, adding extra compile-time checks to the code.

But there were a few that didn't work too:
#define __force           __attribute__((force))
#define __acquires(x)     __attribute__((context(x,0,1)))
#define __releases(x)     __attribute__((context(x,1,0)))
#define __acquire(x)      __context__(x,1)
#define __release(x)      __context__(x,-1)

Actually, __acquires and __releases didn't cause any problems, but __acquire and __release caused a complaint about __context__, and I didn't see any further setup for that. I gather __context__(a,b) adds the second argument to a compiler variable named by the first argument (passed into the macro), while context(a,b,c) says that the compiler variable named by the first argument (passed into the macro) must have value "b" at the start of the function and value "c" when the function returns. But either those definitions are hiding somewhere else, or they're one of the areas where GCC has a feature that LLVM hasn't picked up.

Likewise it didn't seem to like the attribute force, but that just seems to suppress warnings for an odd cast, so it's not so critical for now. I'm way far away from attempting to get all the warnings out of the code!

Anyway, this let me blow another hour or more removing all the #prama pack statements. In return, though, it gave me a lot more confidence that the structs would come out as intended, and made it immensely easier to update the code to the driver included in a newer Linux release, when I later got around to that.

<< Prev: Xcode vs. AppCode       Next: Plugging Gaps >>

Xcode vs. AppCode


Xcode: just plain awful, or actually unusable?

All right, they moved my cheese again. Maybe I'm just not used to it. I'm definitely not used to it.

Some of the things I don't like:
  • Right-clicking to navigate from a name (struct, function, variable, etc.) to where it's defined (using the mouse as opposed to a keystroke)
  • An editor area with only one file at a time (as opposed to tabs), forcing me to navigate a large project with the whole project view instead of quick-switching between a subset of files.
  • The alternative of opening a jillion separate source windows, again with no convenient way to switch to just the one I want
  • The project settings, which is sort of inexplicable and has a monstrous number of options with no description or help
  • The lack of an obvious place to map all project files to source/resources/syntax highlight but don't include in output, etc.
  • The inexplicable use of includes and frameworks. You can #include seemingly anything you like. You can also add frameworks to the project, though it's not clear what that does. It doesn't seem to alter which includes work, and it doesn't seem to put the frameworks into the list of required frameworks in the Info.plist
  • The way the plist editor defaults to an unusable tree view (where among other problems, it takes between one and three clicks on a control to activate the control), and the plain text view is hard to find
  • Lack of an easy way to rename files
  • Odd Git integration, where files you add and edit are sometimes added and sometimes left untracked
  • Strange separation of source code directories and project "groups"
  • Completely inconsistent build results where a project with a trillion errors can emit just one output error, and the if you fix that you get a different one, and then if you fix that you get 60 errors and a message that you've reached the error limit, but then if you fix one of those you get 35 errors and a message that you've still reached the error limit, and etc.

I could probably go on. Every time I use it, it seems like something comes up. And I spend my time hunting around or Googling or endlessly clicking and scrolling instead of doing something useful. I don't mean to just bash the product, but I don't feel productive with it.

On a side note, I once didn't feel productive with IntelliJ either. But when you start IntelliJ, it offers an endless list of tips that ease you over the learning curve. Keyboard shortcuts. Functions otherwise hidden in menus you might not think to look at. Features that make you more productive. Sometimes I just dismiss that dialog without looking, but at least a couple times I've sat there and clicked through it for a while, and I've been rewarded. Where's the "100 ways to be more productive in Xcode" in Xcode?

So then I thought to look over at the JetBrains Web site. I know they've branched out into IDEs for other languages besides Java. Could there be one for C and C++?

There could! Actually, there could be two, which was a bit of a problem because which do you pick? Choosing the more OS X-oriented one, I went with AppCode. When I fired it up and went to create a new project it offered an option to build an IOKit kext, just same as Xcode. That made me think I probably picked the right one.

Well, that might be true, except I quickly found there are two near-crippling bugs in AppCode 3.3 and 3.4 EAP (as of this writing).
  1. If you #include IOKit features such as IOKit/IOLib.h or IOKit/pci/IOPCIDevice.h, AppCode doesn't recognize them. They're shown in red, and code highlighting also shows anything that came from them in red. This has a huge trickle-down effect, as basic data types such as u32 aren't found, and a lot of errors crop up like "incompatible assignment UInt32 to int". It means in a typical source file, there are hundreds of lines highlighted as errors, and the right gutter in the editor window is nearly solid red. Among other things, this makes the scroll knob nearly invisible, depending on the editor color scheme.
  2. When I do a build, I often get on the order of 2 errors and 500 warnings. Unfortunately, the build messages window shows the chain of includes for every file that has a warning or error, often several times. The bottom line is, the build output window has (in the test build I just did) 1665 lines of output, including 2 errors. It shows perhaps 15-20 lines at a time. It normally does not jump right to the error. Using the "don't show warnings" control reduces that to only 908 lines. Somewhere in those 908 lines, of which we'll be generous and say you can see 20, is your error. It might take longer to scroll around and FIND the error than to actually fix it!

The first problem I assume they'll just fix someday.

The second, I'm not so sure. In this one area, it made me really appreciate Xcode, which after a build lists every file with a warning or error exactly once, with a red symbol next to the ones with errors. It's completely obvious where the errors are.

So I'm left with a tough choice.

But at the end of the day, I find that while both IDEs bother me, AppCode bothers me less. Maybe I should go to the trouble of identifying the context-menu option in Xcode to navigate to the definition of a name, and assign a keyboard shortcut to it. Then if I could get Find in Xcode to navigate results as easily as Incremental Search in AppCode, I might stick with Xcode until they fix the syntax highlighting problem in AppCode.

Or, I suppose, I could try to use emacs for more than just text editing. I hear some people use it for code, too.   :)

The Other Shoe Drops...

<< Prev: Parsing Firmware       Next: Xcode vs. AppCode >>

The bad news is, I did more reading. I looked at more code, took more notes, read more background material. I'm becoming increasingly convinced my approach isn't going to work. If I try to rewrite all this code into C++, I'll never finish, and there probably won't even be decent progress (in terms of actually connecting to wireless networks) along the way. There's just too much -- driver code, PCIe interface code, MVM firmware interface code, 802.11 packet-handling code, and etc. I might be able to set up and initialize a card, but I can’t see that I'd actually be able to operate it.

Worse, all the simplifications I thought I'd make don't seem that simple. Rather than just dropping a couple files or functions because, for instance, no Mac has a hardware switch to disable wireless radios, the code for "RFKill" seems to be in little bits all over the place. Taking it out might be harder than leaving it in!

So I think I've been going about this backward. I think what I should be doing is attempting a minimal-changes type port, where I use the existing code to whatever extent is possible. Yes, I'll have to change some things I've already done, OK perhaps a lot of things I've already done, but maybe I'll find some shortcuts. Macros to work around missing Linux APIs or whatever.

In any case, I think that's maybe my only hope of achieving functionality in the not-amazingly-long term. Then once I get it all working, maybe that’s the time to think about migrating more code into C++ and making the organization more pleasing to me.

There will still be plenty of challenges, like the part where it expects the Linux wireless API to do something or other (such as give it connection settings for a network).

So really, the next step is just to sanity-check that option. Is it even reasonable to think I can fill all the gaps between Linux and OS X without a substantial rewrite? If I pull in ALL the iwlwifi driver code, and all the Linux headers they rely on for various structs and constants, can I get back to the working firmware loading-and-parsing code with the new strategy?

And if so, does it seem like a better or worse approach?

<< Prev: Parsing Firmware       Next: Xcode vs. AppCode >>

Parsing Firmware

<< Prev: #include Woes       Next: The Other Shoe Drops >>

After doing a little research and proving that I could load a driver, I figured it was time for some real work. But first, a little background on firmware.

About the Firmware

The firmware for the Intel WiFi cards is distributed as a single file. However, there are at least two major styles of firmware, and each one is broken up into many sections. From a hardware perspective, there's the older models that use DVM firmware vs. the newer models that use MVM firmware. From the firmware file perspective, there's the v1/v2 firmware vs. the TLV firmware. I'm not actually sure whether the DVM/MVM axis is separate from the v1v2/TLV axis.

Since I'm only targeting fairly current hardware and current firmware, I only need to work with MVM hardware and TLV-style firmware. Unfortunately, from all appearances, TLV is the more complex of the two. The large firmware file consists of a header followed by many individual records, each of which has a type and length value and then a variable amount of payload (the size of which is in the length value). In order to be able to load the firmware onto the hardware, you need to parse this file out into all its separate components.

There's a big C function (dominated by an enormous switch statement in a while loop) that handles this. It ends up stashing some values in configuration objects, and copying other sections into freshly-allocated memory that it will hang on to for subsequent restarts or whatever. Then it lets the original file data be freed.

Firmware Parsing Code

So my next step would be calling that function to parse the TLV firmware. This would involve roping a lot of code from the Linux driver into my project. In that code, I found Linux-compiler syntax for byte-alignment in structs, which my OS X compiler was not happy about. It seemed necessary though, because many of the firmware-related structs were packed to map directly to specific bytes in the file, and extra padding would be spoiled that mapping. This effort would test whether I could convert all that successfully, plus hand off from my kext-resource-loading code to the Linux firmware-parsing code.

As I went about pulling in the functions and constants and structs that the firmware parsing used, the project suddenly bloated. I tried to be strategic about what code and files I was bringing into my project, but this referenced that and suddenly I had 10+ files. It wasn't obvious to me why some things were laid out the way they were in the original source, so in some cases I rearranged a bit. I didn't want to try real hard to keep things just as they were in the original driver, because the end goal was to port a lot of it anyway.

For instance, the C code had pretty opaque function tables (structs full of function pointers). In trying to follow the code, I'd land at ops->start(...) but there wasn't any definition of a function "start" anywhere, it was just an entry in this "ops" struct. Then I had to figure out where that was assigned in order to know what the actual function was called in its definition and where that was so I could follow the code. I guess all that makes sense in C, but it seems like an obvious candidate to make a C++ class out of. The C++ code in the other ported drivers I looked at was definitely easier to follow than the original C code.

Bottom line, I copied and rearranged, and soon I had a big pile of code that seemed to be a complete set, but still didn't compile.

Struct Alignment

For whatever reason, it seems that certain hardware is much more efficient at reading and writing values that start on a 32-bit or 64-bit boundary (or other alignments, for less common cases). For instance, on a 64-bit system like OS X, pointers need to be aligned on 64-bit boundaries.

Thus, look at a struct like this on 64-bit OS X:
struct foo {
    UInt32 a;
    some *b;
    SInt16 c;
    some *d;
}
Adding up the field sizes gives 4+8+2+8=22 bytes. But the actual struct takes 4+4(padding)+8+2+6(padding)+8 = 32 bytes. There's padding introduced after "a" to allow "b" to be on a 64-bit boundary, and likewise padding after "c" to allow "d" to be on a 64-bit boundary.

The 64-bit developer docs have an easy suggestion to fix this: just reorder the fields in the struct. If they went b,d,a,c there wouldn't need to be padding.

Well, here's the problem. If a struct with assorted field sizes like that was used to map to a section of a firmware file, I couldn't just reorder the struct without making the fields point to the wrong bytes in the firmware file. And if there isn't padding in the firmware file, there needs to also be no padding in the struct, or again, the fields will point to the wrong bytes in the file.

The Linux driver code makes it work with the __packed keyword like this:
struct iwl_fw_dbg_reg_op {
 u8 op;
 u8 reserved[3];
 __le32 addr;
 __le32 val;
} __packed;

Whereas OS X seems to prefer it with a compiler annotation (pragma) like this:
#pragma pack(1)
struct iwl_fw_dbg_reg_op {
 u8 op;
 u8 reserved[3];
 __le32 addr;
 __le32 val;
};// __packed;
#pragma options align=reset

What made me nervous was the more complicated ones, like the ieee80211.h Linux header file. That one has structs aligned(2), packed structs, and (if you look for struct ieee80211_mgmt) a struct aligned(2) containing packed sub-structs, unions of packed sub-structs, and even a sub-union with packed sub-structs.

I took a swing at converting all that to the corresponding pragma syntax, but I can't say I had any real confidence it would work. You know, what if changing the pragma and back mid-struct didn't work?

Got a better idea? It turns out my entire approach here was flawed. But naturally, I did not discover that until later. We'll come back to struct alignment in a future post.

Memory Allocation

The next problem was that the firmware-parsing logic had at least a little memory allocation going on, and naturally the syntax for that differs as well.

On Linux, it went like this:
pieces = kzalloc(sizeof(*pieces), GFP_KERNEL);
...
kmemdup(pieces->dbg_dest_tlv,
 sizeof(*pieces->dbg_dest_tlv) +
 sizeof(pieces->dbg_dest_tlv->reg_ops[0]) *
 drv->fw.dbg_dest_reg_num, GFP_KERNEL);
...
kfree(pieces);

Well, OS X doesn't have kzalloc, kmemdup, or kfree.

I though at the time kzalloc was the zone allocator. For instance, in OS X, many of the higher-level memory allocation functions pull memory out of "zones" of various sizes (32 bytes, 64 bytes, etc.) and always give you an allocation of that size. Therefor if you ask for 33 bytes, you actually get an allocation of 64 bytes from the next-larger zone.

Well, it turns out I was wrong (more about that, and the corresponding crash, later). But for now, I replaced kzalloc with IOMalloc (which I gather also uses a zone allocator under the covers), and kfree with IOFree. The problem there was that IOFree needs to be passed the original allocation size, so I had to add some fields and logic to track the allocation sizes so that I could use them when the time came around to free. I'm not sure I got that right, so there could be a leak there, but at this point I was mainly aiming for "working" over "working perfectly".

kmemdup was trickier. I found this definition of kmemdup:
void *kmemdup(const void *src, size_t len, gfp_t gfp)
{
    void *p;

    p = kmalloc_track_caller(len, gfp);
    if (p)
        memcpy(p, src, len);
    return p;
}
It looks like it allocates memory and then, if successful, copies the contents of the thing passed to it into that new memory and returns it. Actually, I took the easiest route and just copied the whole function into my project, except I again used IOMalloc instead of the oddball kmalloc_track_caller.

Other Odds and Ends

I did put the firmware parsing logic into a new C++ class. It ended up with a bunch of static utility functions down at the bottom. I could have pulled those into the class and thereby eliminated at least one of the arguments from each one... but I wasn't yet sure whether I wanted to do that. I sort of had it in my mind that I might move them again.

Then I had to comment out all the Linux library imports from the headers I did bring in, and add the linux-porting.h header file that I brought in from another driver port. That handled some things like common constants, macros, and typedefs that people had run into before.

Finally, I commented out a few fields here and there that used data types I wasn't ready to deal with yet. Overall it was a bit of a cleanup operation, but I was trying to avoid any major decisions about how to re-implement things. When in doubt, comment out.

Results

Finally, everything compiled again. I gave it a spin on my test machine, and...

Much to my surprise, the firmware-parsing logic all seemed to pretty much just work! I got a couple debug lines I took to be at least warnings (such as "GSCAN is supported but capabilities TLV is unavailable"), but with a more detailed inspection of the firmware files in a hex editor, it seems to have (and not have) exactly the chunks the parsing code said. Huh.

(Side note: what does it mean when you find success surprising?)

Kext Resource Caching

So then I figured I better try all five firmware files downloadable for the MVM firmware models. I didn't want to go claiming everything was working fine and have it turn out that only one model worked. So I put all the files into the kext Resources and hardcoded the IDs into the driver rather than detecting the real hardware.

Naturally, the first additional model I tried failed. With a little debugging, I found that for some reason my file-loading code refused to load any but the original firmware file.

Eventually I supposed there must be some kind of caching going on so it "remembered" the version of my kext with only the one firmware resource. I rebooted the machine, and surprise, surprise, it all worked. I've since found a passing reference to caching kext resources, but not a full explanation.

From "man kextutil":
-c, -no-caches
  Ignore any repository cache files and scan all kext bundles to
  gather information.  If this option is not given, kextutil
  attempts to use cache files and (when running as root) to create
  them if they are out of date or don't exist.

Sample Code and Output

This version of the code is available here (and a build here). However, it may or may not work for you. The kzalloc thing means that sometimes the counters come out totally bogus, and it seems like a crash follows quickly thereafter.

<< Prev: #include Woes       Next: The Other Shoe Drops >>

Tuesday, March 22, 2016

#include Woes

<< Prev: Loading Firmware       Next: Parsing Firmware >>

Is Apple somehow unaware that in order to use their lovely APIs, you have to #include them first? It makes the IDE (not to mention the compiler) a little cranky when you omit that part.

Yet, have a look at this documentation for OSKextRequestResource again. It says it's documenting OSKextLib.h, so you might try something like this:

#include <OSKextLib.h>

Nope. That header is itself in something, or under something. In or under what? No indication. Maybe Apple has some other documentation on how to use that? Perhaps under Creating a Device Driver with Xcode?

Well, that gives a code sample that actually includes a #include (yay!), but it's not the one we need here.

This kind of problem came up a number of times. I needed to use an OSData object to store the results ultimately generated by the OSKextResourceRequest. Let's look at the OSData Class Reference. That page doesn't even name the header file it's in.

There's also a detailed article on kernel data structures including OSData, with code samples, big and small. Neither the big ones nor the small ones include the #includes. As far as I can tell the article never mentions what to #include either. Maybe if you always write tidy, one-class drivers that do virtually nothing, it just works based on what a "heavyweight" header such as #include <IOKit/IOLib.h> ropes in, but that's not that helpful in general.

So, it comes down to the source code I guess. Googling OSKextRequestResource turns up this source code. From the URL, you might decide to try

#include <libkern/OSKextLib.h>
or
#include <libkern/libkern/OSKextLib.h>

That first one turns out to work.

There’s also the good-old-grep method:

cd /Applications/Xcode.app/Contents/Developer/
cd Platforms/MacOSX.platform/Developer/SDKs/
cd MacOSX10.11.sdk/System/Library/Frameworks/
grep -r "OSKextRequestResource" *

The output is a bit of a mess, but you'll find Kernel.framework/Headers/libkern/OSKextLib.h in there, suggesting the libkern/OSKextLib.h used above.

For what it's worth, OSData seems to be defined in libkern/c++/OSData.h (try searching for "class OSData" if you don’t want to be overwhelmed by occurrences of just "OSData"), though again it tends to get roped in by the IOLib imports such as IOService.h. I have to think there's some middlin' header that includes just the data structures like OSData and OSDictionary and stuff, but I haven't come across it yet.

Why is this so difficult?

It makes me think that one of the huge advantages of Java is the documentation. Which in turn, is build on the package system. If you look at HashMap, for instance, right above the big words Class HashMap are the little words java.util and immediately you know precisely the import you need to write to actually use this class.

Even supposing C/C++ isn't amenable to that level of automated quality documentation generation (though why should it not be?), Apple certainly had the opportunity in the articles like this one to make the needed #include clear.

Oh, well. Grep on.

<< Prev: Loading Firmware       Next: Parsing Firmware >>

Loading Firmware

<< Prev: Basic Proof of Concept       Next: #include Woes >>

Loading Firmware

After recognizing the hardware, the next thing I wanted to do was identify and load a firmware file that would work for the hardware in question. The code with the device IDs that I originally found (in here) included a configuration struct that included most of the firmware file name -- everything except the version number -- and the range of acceptable version numbers (see, for instance, all the config objects in the bottom half of this file).

It seems like the Linux API lets you request a firmware file and get a callback with either the file data or an error (request_firmware_nowait). So the iwlwifi driver just attempts to load all possible firmware files from highest version to lowest version until it finds a match. To resolve that kind of request, Linux talks to a user-space process that looks for the files in question in a directory such as /lib/firmware, and the user is expected to put the firmware files there.

I didn't see any similar infrastructure in OS X -- something to sit on top of data or configuration files in a directory and serve them up to kernel extensions upon request. Long-term, this might be a challenge. It would be obnoxious to build a user-space daemon to sit on a directory like /lib/firmware and the methods for the kernel to communicate with it. Plus there would need to be a way to ensure that loaded before any driver that needed to request firmware from it. But other drivers (even existing Bluetooth drivers) have the same problem, and without some common option like that every driver would have to implement its own solution, which seems unfortunate.

Still, there's a short-term solution. Kexts can include arbitrary files in their Resources directory, and I found the OSKextRequestResource call load file data out of Resources. So I could load the firmware file from there. I tried bundling a firmware file into my kext and tagging it as a build resource and loading it with a hardcoded file name, and that all worked! (See my original loadFirmwareSync method, and the callback just above it.)

Actual firmware naming logic

The next step was pulling in the logic that actually held a specific firmware file name for each piece of hardware. It turned out there were five possible firmware files. I briefly considered just building my own mapping of device IDs to file names, but all the information was already in the Linux driver code... I just had to find a way to take advantage of it.

The problems were thus:
  • The code was split across three files. One defined all the 7000-series device configuration possibilities. One defined the 8000-series device configuration possibilities. And the third listed every possible device by ID, along with which configuration it mapped to. Often many specific devices mapped to a particular configuration, and more than one configuration mapped to a given firmware file.
  • All three files were .c files instead of .h files, meaning it wasn't a completely straightforward #include to pull that logic in
  • The third file had a bunch of additional logic. It was all related to how the driver registered to be loaded when those particular devices were present. The more I looked the more it seemed like the same kind of plumbing I had built into my Kext -- the Linux equivalent of the IOKit registration, startup, and shutdown processes. I didn't need that code.
So I took the easy route of converting the first two files to headers (iwl-7000.h and iwl-8000.h), and pulling the array of devices out of the third into a new header file (device-list.h). You may notice a few changes. In the first two, I just commented out the calls that link modules to firmware files, since that's all Linux-specific plumbing. If you compare the third to the device list in the original, you'll see that I simplified the struct that holds the device information. I didn't need the full Linux pci_device_id struct, just the device ID, subdevice ID, and configuration struct.

Then, I duplicated the logic of the Linux driver that checked available versions of the firmware file for each device. To do that, I sort of adopted some of the structs used by the Linux code, such as iwl_drv, iwl_cfg, and iwl_trans. These I pulled out of wherever they were defined into my own code. So I ended up with a revised loadFirmwareSync method and a few header files that I either migrated directly (like iwl-config.h) or created to hold a grab bag of #defines and structs from files I wasn't ready to convert.

The Joys of Kext Development

It was about this time I found out what I was messing with. I had allocated some object or other that I freed when I finished using it. Then (whoops) I freed it again in the overall Free function of my driver.

So I tested my new hardware-matching, firmware-loading kext: "sudo kextload AppleIntelWiFIMVM.kext". All good. Yay!

Then I cleaned it up so I'd be able to try a new version later. "sudo kextunload AppleIntelWiFiMVM.kext".

Then my screen turned off.

Then I heard the little Apple chime.

Then my laptop informed me that it had restarted because of a problem.

So much for the "good old days" of Segmentation Fault; Core Dump. Instead I had an inscrutable panic report, with some elaborate complaint about 0xDEADBEEF. I guess this is what passes for humor among kernel developers? Though, I guess I should thank them for bothering to make my pointer point to something obvious when I freed it the first time, so they could tell when I freed it a second time.

Anyway, it appears they weren't kidding around when they said in case of error, the kernel would shut itself down rather than risk filesystem corruption and associated problems. Still, I had like thirty windows open!

Well, I've learned my lesson and now test my builds on another machine. On the upside, I can test on a machine with the actual Intel wireless hardware, and I was able to turn SIP back on for my laptop. But now it's a two-machine operation. I guess that day was coming anyway, since there's only so far to go without using the real hardware.

<< Prev: Basic Proof of Concept       Next: #include Woes >>